AIRGuard: Guarding Agent Actions with Runtime Authority Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Suliu, Zhuang, Haomin, Zhou, Yujun, Han, Yufei, Zhang, Xiangliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910267280982016
author Qin, Suliu
Zhuang, Haomin
Zhou, Yujun
Han, Yufei
Zhang, Xiangliang
author_facet Qin, Suliu
Zhuang, Haomin
Zhou, Yujun
Han, Yufei
Zhang, Xiangliang
contents Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before actions execute. On AgentTrap, AIRGuard reduces Sonnet 4.6 attack success from 36.3% without defense to 5.5%. On DTAP-150, AIRGuard preserves 76.0% benign utility with Haiku 4.5, compared with 52.0% for ARGUS and 42.0% for MELON. An ablation further shows that prompt-only policy helps only modestly, whereas a dedicated runtime authority-control layer gives the agent system direct control over tool-mediated side effects. Code and data are available at https://github.com/Sophie508/AIRGuard.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AIRGuard: Guarding Agent Actions with Runtime Authority Control
Qin, Suliu
Zhuang, Haomin
Zhou, Yujun
Han, Yufei
Zhang, Xiangliang
Cryptography and Security
Artificial Intelligence
Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model Context Protocol tools. This makes agent attacks different from jailbreaks. The harmful step is often not an obviously forbidden output, but an ordinary executable action that becomes unsafe because attacker-controlled context steers authorized access against the user's interest. We identify this failure mode as authority confusion: untrusted resources may inform reasoning, but they must not authorize side effects. We present AIRGuard, a runtime guard that operationalizes least privilege as action-time authorization. AIRGuard normalizes heterogeneous tool calls, derives task authority into step-level authority, tracks source and target trust, simulates sensitive side effects, audits cross-step risk, and enforces decisions before actions execute. On AgentTrap, AIRGuard reduces Sonnet 4.6 attack success from 36.3% without defense to 5.5%. On DTAP-150, AIRGuard preserves 76.0% benign utility with Haiku 4.5, compared with 52.0% for ARGUS and 42.0% for MELON. An ablation further shows that prompt-only policy helps only modestly, whereas a dedicated runtime authority-control layer gives the agent system direct control over tool-mediated side effects. Code and data are available at https://github.com/Sophie508/AIRGuard.
title AIRGuard: Guarding Agent Actions with Runtime Authority Control
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.28914