AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
Fuente:
arXiv
Saved in:
| Main Author: | Yang, Chenglin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026)
by: Aravind, Ashwin
Published: (2026)
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026)
by: Zhuang, Haomin, et al.
Published: (2026)
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
by: Chen, Jizhou, et al.
Published: (2025)
by: Chen, Jizhou, et al.
Published: (2025)
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026)
by: Chia-Pei, et al.
Published: (2026)
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
by: Zhao, Wei, et al.
Published: (2026)
by: Zhao, Wei, et al.
Published: (2026)
Securing GenAI Multi-Agent Systems Against Tool Squatting: A Zero Trust Registry-Based Approach
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Trusted AI Agents in the Cloud
by: Bodea, Teofil, et al.
Published: (2025)
by: Bodea, Teofil, et al.
Published: (2025)
AIRGuard: Guarding Agent Actions with Runtime Authority Control
by: Qin, Suliu, et al.
Published: (2026)
by: Qin, Suliu, et al.
Published: (2026)
Safeguarding AI Agents: Developing and Analyzing Safety Architectures
by: Domkundwar, Ishaan, et al.
Published: (2024)
by: Domkundwar, Ishaan, et al.
Published: (2024)
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare
by: Maiti, Saikat
Published: (2026)
by: Maiti, Saikat
Published: (2026)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
by: Ding, Xuwei, et al.
Published: (2026)
by: Ding, Xuwei, et al.
Published: (2026)
Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability
by: Mo, Jiayun, et al.
Published: (2025)
by: Mo, Jiayun, et al.
Published: (2025)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents
by: Uchibeke, Uchi
Published: (2026)
by: Uchibeke, Uchi
Published: (2026)
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
by: Jin, Xisen, et al.
Published: (2026)
by: Jin, Xisen, et al.
Published: (2026)
ToolTweak: An Attack on Tool Selection in LLM-based Agents
by: Sneh, Jonathan, et al.
Published: (2025)
by: Sneh, Jonathan, et al.
Published: (2025)
OpenPort Protocol: A Security Governance Specification for AI Agent Tool Access
by: Zhu, Genliang, et al.
Published: (2026)
by: Zhu, Genliang, et al.
Published: (2026)
Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use
by: Zhang, Wuyang, et al.
Published: (2026)
by: Zhang, Wuyang, et al.
Published: (2026)
Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives
by: Son, Daeyeon
Published: (2026)
by: Son, Daeyeon
Published: (2026)
SoK: Trust-Authorization Mismatch in LLM Agent Interactions
by: Shi, Guanquan, et al.
Published: (2025)
by: Shi, Guanquan, et al.
Published: (2025)
Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments
by: Goel, Hardik
Published: (2026)
by: Goel, Hardik
Published: (2026)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
Autonomous Action Runtime Management(AARM):A System Specification for Securing AI-Driven Actions at Runtime
by: Errico, Herman
Published: (2026)
by: Errico, Herman
Published: (2026)
Agent Safety Alignment via Reinforcement Learning
by: Sha, Zeyang, et al.
Published: (2025)
by: Sha, Zeyang, et al.
Published: (2025)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents
by: Aonzo, Simone, et al.
Published: (2026)
by: Aonzo, Simone, et al.
Published: (2026)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
Mind the Web: The Security of Web Use Agents
by: Shapira, Avishag, et al.
Published: (2025)
by: Shapira, Avishag, et al.
Published: (2025)
Measuring Safety Alignment Effects in Autonomous Security Agents
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
by: Zhang, Yixiang, et al.
Published: (2026)
by: Zhang, Yixiang, et al.
Published: (2026)
AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
by: Zheng, Ye, et al.
Published: (2025)
by: Zheng, Ye, et al.
Published: (2025)
Zero-Trust Runtime Verification for Agentic Payment Protocols: Mitigating Replay and Context-Binding Failures in AP2
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs
by: Yang, Ya-Ting, et al.
Published: (2026)
by: Yang, Ya-Ting, et al.
Published: (2026)
A2AS: Agentic AI Runtime Security and Self-Defense
by: Neelou, Eugene, et al.
Published: (2025)
by: Neelou, Eugene, et al.
Published: (2025)
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
by: Yang, Zhi, et al.
Published: (2026)
by: Yang, Zhi, et al.
Published: (2026)
ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
by: Jiang, Xiaochong, et al.
Published: (2026)
by: Jiang, Xiaochong, et al.
Published: (2026)
Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space
by: Zhou, Xingfu, et al.
Published: (2025)
by: Zhou, Xingfu, et al.
Published: (2025)
Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes
by: Metere, Alfredo
Published: (2026)
by: Metere, Alfredo
Published: (2026)
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
by: Motwani, Sumeet Ramesh, et al.
Published: (2024)
Similar Items
-
AgentWall: A Runtime Safety Layer for Local AI Agents
by: Aravind, Ashwin
Published: (2026) -
AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills
by: Zhuang, Haomin, et al.
Published: (2026) -
AgentGuard: Repurposing Agentic Orchestrator for Safety Evaluation of Tool Orchestration
by: Chen, Jizhou, et al.
Published: (2025) -
IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection
by: Chia-Pei, et al.
Published: (2026) -
ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection
by: Zhao, Wei, et al.
Published: (2026)