Agent Safety Alignment via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sha, Zeyang, Tian, Hanling, Xu, Zhuoer, Cui, Shiwen, Meng, Changhua, Wang, Weiqiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912476117860352
author Sha, Zeyang
Tian, Hanling
Xu, Zhuoer
Cui, Shiwen
Meng, Changhua
Wang, Weiqiang
author_facet Sha, Zeyang
Tian, Hanling
Xu, Zhuoer
Cui, Shiwen
Meng, Changhua
Wang, Weiqiang
contents The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents, empowered to execute external functions, are vulnerable to both user-initiated threats (e.g., adversarial prompts) and tool-initiated threats (e.g., malicious outputs from compromised tools). In this paper, we propose the first unified safety-alignment framework for tool-using agents, enabling models to handle both channels of threat via structured reasoning and sandboxed reinforcement learning. We introduce a tri-modal taxonomy, including benign, malicious, and sensitive for both user prompts and tool responses, and define a policy-driven decision model. Our framework employs a custom-designed sandbox environment that simulates real-world tool execution and allows fine-grained reward shaping. Through extensive evaluations on public and self-built benchmarks, including Agent SafetyBench, InjecAgent, and BFCL, we demonstrate that our safety-aligned agents significantly improve resistance to security threats while preserving strong utility on benign tasks. Our results show that safety and effectiveness can be jointly optimized, laying the groundwork for trustworthy deployment of autonomous LLM agents.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agent Safety Alignment via Reinforcement Learning
Sha, Zeyang
Tian, Hanling
Xu, Zhuoer
Cui, Shiwen
Meng, Changhua
Wang, Weiqiang
Artificial Intelligence
Cryptography and Security
The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents, empowered to execute external functions, are vulnerable to both user-initiated threats (e.g., adversarial prompts) and tool-initiated threats (e.g., malicious outputs from compromised tools). In this paper, we propose the first unified safety-alignment framework for tool-using agents, enabling models to handle both channels of threat via structured reasoning and sandboxed reinforcement learning. We introduce a tri-modal taxonomy, including benign, malicious, and sensitive for both user prompts and tool responses, and define a policy-driven decision model. Our framework employs a custom-designed sandbox environment that simulates real-world tool execution and allows fine-grained reward shaping. Through extensive evaluations on public and self-built benchmarks, including Agent SafetyBench, InjecAgent, and BFCL, we demonstrate that our safety-aligned agents significantly improve resistance to security threats while preserving strong utility on benign tasks. Our results show that safety and effectiveness can be jointly optimized, laying the groundwork for trustworthy deployment of autonomous LLM agents.
title Agent Safety Alignment via Reinforcement Learning
topic Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2507.08270