Building a Foundational Guardrail for General Agentic Systems via Synthetic Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yue, Hua, Hang, Zhou, Yujun, Jing, Pengcheng, Nagireddy, Manish, Padhi, Inkit, Dolcetti, Greta, Xu, Zhangchen, Chaudhury, Subhajit, Rawat, Ambrish, Nedoshivina, Liubov, Chen, Pin-Yu, Sattigeri, Prasanna, Zhang, Xiangliang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912642370633728
author Huang, Yue
Hua, Hang
Zhou, Yujun
Jing, Pengcheng
Nagireddy, Manish
Padhi, Inkit
Dolcetti, Greta
Xu, Zhangchen
Chaudhury, Subhajit
Rawat, Ambrish
Nedoshivina, Liubov
Chen, Pin-Yu
Sattigeri, Prasanna
Zhang, Xiangliang
author_facet Huang, Yue
Hua, Hang
Zhou, Yujun
Jing, Pengcheng
Nagireddy, Manish
Padhi, Inkit
Dolcetti, Greta
Xu, Zhangchen
Chaudhury, Subhajit
Rawat, Ambrish
Nedoshivina, Liubov
Chen, Pin-Yu
Sattigeri, Prasanna
Zhang, Xiangliang
contents While LLM agents can plan multi-step tasks, intervening at the planning stage-before any action is executed-is often the safest way to prevent harm, since certain risks can lead to severe consequences once carried out. However, existing guardrails mostly operate post-execution, which is difficult to scale and leaves little room for controllable supervision at the plan level. To address this challenge, we highlight three critical gaps in current research: data gap, model gap, and evaluation gap. To close the data gap, we introduce AuraGen, a controllable engine that (i) synthesizes benign trajectories, (ii) injects category-labeled risks with calibrated difficulty, and (iii) filters outputs via an automated reward model, producing large and reliable corpora for pre-execution safety. To close the guardian model gap, we propose a foundational guardrail Safiron, combining a cross-planner adapter with a compact guardian model. The adapter unifies different input formats, while Safiron flags risky cases, assigns risk types, and generates rationales; trained in two stages with a broadly explored data recipe, Safiron achieves robust transfer across settings. To close the evaluation gap, we release Pre-Exec Bench, a realistic benchmark covering diverse tools and branching trajectories, which measures detection, fine-grained categorization, explanation, and cross-planner generalization in human-verified scenarios. Extensive experiments demonstrate consistent gains of the proposed guardrail over strong baselines on Pre-Exec Bench, and ablations further distill actionable practices, providing a practical template for safer agentic systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_09781
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
Huang, Yue
Hua, Hang
Zhou, Yujun
Jing, Pengcheng
Nagireddy, Manish
Padhi, Inkit
Dolcetti, Greta
Xu, Zhangchen
Chaudhury, Subhajit
Rawat, Ambrish
Nedoshivina, Liubov
Chen, Pin-Yu
Sattigeri, Prasanna
Zhang, Xiangliang
Machine Learning
Artificial Intelligence
Computation and Language
While LLM agents can plan multi-step tasks, intervening at the planning stage-before any action is executed-is often the safest way to prevent harm, since certain risks can lead to severe consequences once carried out. However, existing guardrails mostly operate post-execution, which is difficult to scale and leaves little room for controllable supervision at the plan level. To address this challenge, we highlight three critical gaps in current research: data gap, model gap, and evaluation gap. To close the data gap, we introduce AuraGen, a controllable engine that (i) synthesizes benign trajectories, (ii) injects category-labeled risks with calibrated difficulty, and (iii) filters outputs via an automated reward model, producing large and reliable corpora for pre-execution safety. To close the guardian model gap, we propose a foundational guardrail Safiron, combining a cross-planner adapter with a compact guardian model. The adapter unifies different input formats, while Safiron flags risky cases, assigns risk types, and generates rationales; trained in two stages with a broadly explored data recipe, Safiron achieves robust transfer across settings. To close the evaluation gap, we release Pre-Exec Bench, a realistic benchmark covering diverse tools and branching trajectories, which measures detection, fine-grained categorization, explanation, and cross-planner generalization in human-verified scenarios. Extensive experiments demonstrate consistent gains of the proposed guardrail over strong baselines on Pre-Exec Bench, and ablations further distill actionable practices, providing a practical template for safer agentic systems.
title Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.09781