SafeAgent: A Runtime Protection Architecture for Agentic Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Hailin, Ilyushin, Eugene, Ni, Jie, Zhu, Min
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913045468413952
author Liu, Hailin
Ilyushin, Eugene
Ni, Jie
Zhu, Min
author_facet Liu, Hailin
Ilyushin, Eugene
Ni, Jie
Zhu, Min
contents Large language model (LLM) agents are vulnerable to prompt-injection attacks that propagate through multi-step workflows, tool interactions, and persistent context, making input-output filtering alone insufficient for reliable protection. This paper presents SafeAgent, a runtime security architecture that treats agent safety as a stateful decision problem over evolving interaction trajectories. The proposed design separates execution governance from semantic risk reasoning through two coordinated components: a runtime controller that mediates actions around the agent loop and a context-aware decision core that operates over persistent session state. The core is formalized as a context-aware advanced machine intelligence and instantiated through operators for risk encoding, utility-cost evaluation, consequence modeling, policy arbitration, and state synchronization. Experiments on Agent Security Bench (ASB) and InjecAgent show that SafeAgent consistently improves robustness over baseline and text-level guardrail methods while maintaining competitive benign-task performance. Ablation studies further show that recovery confidence and policy weighting determine distinct safety-utility operating points.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17562
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SafeAgent: A Runtime Protection Architecture for Agentic Systems
Liu, Hailin
Ilyushin, Eugene
Ni, Jie
Zhu, Min
Artificial Intelligence
Multiagent Systems
Large language model (LLM) agents are vulnerable to prompt-injection attacks that propagate through multi-step workflows, tool interactions, and persistent context, making input-output filtering alone insufficient for reliable protection. This paper presents SafeAgent, a runtime security architecture that treats agent safety as a stateful decision problem over evolving interaction trajectories. The proposed design separates execution governance from semantic risk reasoning through two coordinated components: a runtime controller that mediates actions around the agent loop and a context-aware decision core that operates over persistent session state. The core is formalized as a context-aware advanced machine intelligence and instantiated through operators for risk encoding, utility-cost evaluation, consequence modeling, policy arbitration, and state synchronization. Experiments on Agent Security Bench (ASB) and InjecAgent show that SafeAgent consistently improves robustness over baseline and text-level guardrail methods while maintaining competitive benign-task performance. Ablation studies further show that recovery confidence and policy weighting determine distinct safety-utility operating points.
title SafeAgent: A Runtime Protection Architecture for Agentic Systems
topic Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2604.17562