SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xueyang, Wang, Weidong, Lu, Lin, Shi, Jiawen, Tie, Guiyao, Xu, Yongtian, Chen, Lixing, Zhou, Pan, Gong, Neil Zhenqiang, Sun, Lichao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909694536187904
author Zhou, Xueyang
Wang, Weidong
Lu, Lin
Shi, Jiawen
Tie, Guiyao
Xu, Yongtian
Chen, Lixing
Zhou, Pan
Gong, Neil Zhenqiang
Sun, Lichao
author_facet Zhou, Xueyang
Wang, Weidong
Lu, Lin
Shi, Jiawen
Tie, Guiyao
Xu, Yongtian
Chen, Lixing
Zhou, Pan
Gong, Neil Zhenqiang
Sun, Lichao
contents Large Language Model (LLM)-based agents are increasingly deployed in real-world applications such as "digital assistants, autonomous customer service, and decision-support systems", where their ability to "interact in multi-turn, tool-augmented environments" makes them indispensable. However, ensuring the safety of these agents remains a significant challenge due to the diverse and complex risks arising from dynamic user interactions, external tool usage, and the potential for unintended harmful behaviors. To address this critical issue, we propose AutoSafe, the first framework that systematically enhances agent safety through fully automated synthetic data generation. Concretely, 1) we introduce an open and extensible threat model, OTS, which formalizes how unsafe behaviors emerge from the interplay of user instructions, interaction contexts, and agent actions. This enables precise modeling of safety risks across diverse scenarios. 2) we develop a fully automated data generation pipeline that simulates unsafe user behaviors, applies self-reflective reasoning to generate safe responses, and constructs a large-scale, diverse, and high-quality safety training dataset-eliminating the need for hazardous real-world data collection. To evaluate the effectiveness of our framework, we design comprehensive experiments on both synthetic and real-world safety benchmarks. Results demonstrate that AutoSafe boosts safety scores by 45% on average and achieves a 28.91% improvement on real-world tasks, validating the generalization ability of our learned safety strategies. These results highlight the practical advancement and scalability of AutoSafe in building safer LLM-based agents for real-world deployment. We have released the project page at https://auto-safe.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17735
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
Zhou, Xueyang
Wang, Weidong
Lu, Lin
Shi, Jiawen
Tie, Guiyao
Xu, Yongtian
Chen, Lixing
Zhou, Pan
Gong, Neil Zhenqiang
Sun, Lichao
Artificial Intelligence
68T07
I.2.6
Large Language Model (LLM)-based agents are increasingly deployed in real-world applications such as "digital assistants, autonomous customer service, and decision-support systems", where their ability to "interact in multi-turn, tool-augmented environments" makes them indispensable. However, ensuring the safety of these agents remains a significant challenge due to the diverse and complex risks arising from dynamic user interactions, external tool usage, and the potential for unintended harmful behaviors. To address this critical issue, we propose AutoSafe, the first framework that systematically enhances agent safety through fully automated synthetic data generation. Concretely, 1) we introduce an open and extensible threat model, OTS, which formalizes how unsafe behaviors emerge from the interplay of user instructions, interaction contexts, and agent actions. This enables precise modeling of safety risks across diverse scenarios. 2) we develop a fully automated data generation pipeline that simulates unsafe user behaviors, applies self-reflective reasoning to generate safe responses, and constructs a large-scale, diverse, and high-quality safety training dataset-eliminating the need for hazardous real-world data collection. To evaluate the effectiveness of our framework, we design comprehensive experiments on both synthetic and real-world safety benchmarks. Results demonstrate that AutoSafe boosts safety scores by 45% on average and achieves a 28.91% improvement on real-world tasks, validating the generalization ability of our learned safety strategies. These results highlight the practical advancement and scalability of AutoSafe in building safer LLM-based agents for real-world deployment. We have released the project page at https://auto-safe.github.io/.
title SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
topic Artificial Intelligence
68T07
I.2.6
url https://arxiv.org/abs/2505.17735