Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Bazinska, Julia, Mathys, Max, Casucci, Francesco, Rojas-Carulla, Mateo, Davies, Xander, Souly, Alexandra, Pfister, Niklas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UK AISI Alignment Evaluation Case-Study
by: Souly, Alexandra, et al.
Published: (2026)
by: Souly, Alexandra, et al.
Published: (2026)
Gandalf the Red: Adaptive Security for LLMs
by: Pfister, Niklas, et al.
Published: (2025)
by: Pfister, Niklas, et al.
Published: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
by: Saha, Shoumik, et al.
Published: (2025)
by: Saha, Shoumik, et al.
Published: (2025)
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024)
by: He, Yifeng, et al.
Published: (2024)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
by: Zhang, Yixiang, et al.
Published: (2026)
by: Zhang, Yixiang, et al.
Published: (2026)
Securing AI Agents with Information-Flow Control
by: Costa, Manuel, et al.
Published: (2025)
by: Costa, Manuel, et al.
Published: (2025)
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
EVMbench: Evaluating AI Agents on Smart Contract Security
by: Wang, Justin, et al.
Published: (2026)
by: Wang, Justin, et al.
Published: (2026)
ClawLess: A Security Model of AI Agents
by: Lu, Hongyi, et al.
Published: (2026)
by: Lu, Hongyi, et al.
Published: (2026)
SoK: Security and Privacy of AI Agents for Blockchain
by: Romandini, Nicolò, et al.
Published: (2025)
by: Romandini, Nicolò, et al.
Published: (2025)
Securing AI Agents Against Prompt Injection Attacks
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
by: Ramakrishnan, Badrinath, et al.
Published: (2025)
AVISE: Framework for Evaluating the Security of AI Systems
by: Lempinen, Mikko, et al.
Published: (2026)
by: Lempinen, Mikko, et al.
Published: (2026)
The Dark Side of LLMs: Agent-based Attack Vectors for System-level Compromise
by: Lupinacci, Matteo, et al.
Published: (2025)
by: Lupinacci, Matteo, et al.
Published: (2025)
Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs
by: Davies, Xander, et al.
Published: (2025)
by: Davies, Xander, et al.
Published: (2025)
A Security Analysis of the OpenClaw AI Agent Framework
by: Suwansathit, Surada, et al.
Published: (2026)
by: Suwansathit, Surada, et al.
Published: (2026)
Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents
by: de Witt, Christian Schroeder, et al.
Published: (2025)
by: de Witt, Christian Schroeder, et al.
Published: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents
by: Zhang, Xiaolei, et al.
Published: (2026)
by: Zhang, Xiaolei, et al.
Published: (2026)
Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents
by: Goswami, Abhishek
Published: (2025)
by: Goswami, Abhishek
Published: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
The End of Trust: How Agentic AI Breaks Security Assumptions
by: Zafar, Osama, et al.
Published: (2026)
by: Zafar, Osama, et al.
Published: (2026)
On the Importance of Backbone to the Adversarial Robustness of Object Detectors
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Provably Secure Agent Guardrail
by: Wu, Benlong, et al.
Published: (2026)
by: Wu, Benlong, et al.
Published: (2026)
ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code
by: Filho, Elzo Brito dos Santos
Published: (2026)
by: Filho, Elzo Brito dos Santos
Published: (2026)
Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents
by: Maloyan, Narek, et al.
Published: (2026)
by: Maloyan, Narek, et al.
Published: (2026)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare
by: Maiti, Saikat
Published: (2026)
by: Maiti, Saikat
Published: (2026)
Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability
by: Mo, Jiayun, et al.
Published: (2025)
by: Mo, Jiayun, et al.
Published: (2025)
Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents
by: Li, Jiaqi, et al.
Published: (2026)
by: Li, Jiaqi, et al.
Published: (2026)
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
by: Cheng, Hao, et al.
Published: (2026)
by: Cheng, Hao, et al.
Published: (2026)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
by: Wu, Xinyi, et al.
Published: (2026)
by: Wu, Xinyi, et al.
Published: (2026)
Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
by: Puppala, Sai, et al.
Published: (2026)
by: Puppala, Sai, et al.
Published: (2026)
Agent Audit: A Security Analysis System for LLM Agent Applications
by: Zhang, Haiyue, et al.
Published: (2026)
by: Zhang, Haiyue, et al.
Published: (2026)
The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates
by: Willis, John M.
Published: (2026)
by: Willis, John M.
Published: (2026)
Agent Security is a Systems Problem
by: Christodorescu, Mihai, et al.
Published: (2026)
by: Christodorescu, Mihai, et al.
Published: (2026)
Security of Internet of Agents: Attacks and Countermeasures
by: Wang, Yuntao, et al.
Published: (2025)
by: Wang, Yuntao, et al.
Published: (2025)
OpenPort Protocol: A Security Governance Specification for AI Agent Tool Access
by: Zhu, Genliang, et al.
Published: (2026)
by: Zhu, Genliang, et al.
Published: (2026)
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
by: Deng, Zehang, et al.
Published: (2024)
by: Deng, Zehang, et al.
Published: (2024)
Similar Items
-
UK AISI Alignment Evaluation Case-Study
by: Souly, Alexandra, et al.
Published: (2026) -
Gandalf the Red: Adaptive Security for LLMs
by: Pfister, Niklas, et al.
Published: (2025) -
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
by: Saha, Shoumik, et al.
Published: (2025) -
Security of AI Agents
by: He, Yifeng, et al.
Published: (2024) -
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)