Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Kellen Tan, Gentile, Anna Lisa, DeLuca, Chad, Ren, Guang-Jie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909752761516032
author Cheng, Kellen Tan
Gentile, Anna Lisa
DeLuca, Chad
Ren, Guang-Jie
author_facet Cheng, Kellen Tan
Gentile, Anna Lisa
DeLuca, Chad
Ren, Guang-Jie
contents The pervasiveness of large language models (LLMs) in enterprise settings has also brought forth a significant amount of risks associated with their usage. Guardrails technologies aim to mitigate this risk by filtering LLMs' input/output text through various detectors. However, developing and maintaining robust detectors faces many challenges, one of which is the difficulty in acquiring production-quality labeled data on real LLM outputs prior to deployment. In this work, we propose backprompting, a simple yet intuitive solution to generate production-like labeled data for health advice guardrails development. Furthermore, we pair our backprompting method with a sparse human-in-the-loop clustering technique to label the generated data. Our aim is to construct a parallel corpus roughly representative of the original dataset yet resembling real LLM output. We then infuse existing datasets with our synthetic examples to produce robust training data for our detector. We test our technique in one of the most difficult and nuanced guardrails: the identification of health advice in LLM output, and demonstrate improvement versus other solutions. Our detector is able to outperform GPT-4o by up to 3.73%, despite having 400x less parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2508_18384
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
Cheng, Kellen Tan
Gentile, Anna Lisa
DeLuca, Chad
Ren, Guang-Jie
Computation and Language
Artificial Intelligence
The pervasiveness of large language models (LLMs) in enterprise settings has also brought forth a significant amount of risks associated with their usage. Guardrails technologies aim to mitigate this risk by filtering LLMs' input/output text through various detectors. However, developing and maintaining robust detectors faces many challenges, one of which is the difficulty in acquiring production-quality labeled data on real LLM outputs prior to deployment. In this work, we propose backprompting, a simple yet intuitive solution to generate production-like labeled data for health advice guardrails development. Furthermore, we pair our backprompting method with a sparse human-in-the-loop clustering technique to label the generated data. Our aim is to construct a parallel corpus roughly representative of the original dataset yet resembling real LLM output. We then infuse existing datasets with our synthetic examples to produce robust training data for our detector. We test our technique in one of the most difficult and nuanced guardrails: the identification of health advice in LLM output, and demonstrate improvement versus other solutions. Our detector is able to outperform GPT-4o by up to 3.73%, despite having 400x less parameters.
title Backprompting: Leveraging Synthetic Production Data for Health Advice Guardrails
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.18384