Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdelnabi, Sahar, Hicks, Chris, Rieck, Konrad, Sadeghi, Ahmad-Reza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
von: Salem, Ahmed, et al.
Veröffentlicht: (2026)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
von: Emerson, Harry, et al.
Veröffentlicht: (2024)
von: Emerson, Harry, et al.
Veröffentlicht: (2024)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
von: Mavali, Sina, et al.
Veröffentlicht: (2026)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)
SkillTester: Benchmarking Utility and Security of Agent Skills
von: Wang, Leye, et al.
Veröffentlicht: (2026)
von: Wang, Leye, et al.
Veröffentlicht: (2026)
Measuring Safety Alignment Effects in Autonomous Security Agents
von: David, Isaac, et al.
Veröffentlicht: (2026)
von: David, Isaac, et al.
Veröffentlicht: (2026)
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
von: Gomaa, Amr, et al.
Veröffentlicht: (2025)
Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems
von: Sharma, Gauri, et al.
Veröffentlicht: (2025)
von: Sharma, Gauri, et al.
Veröffentlicht: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
Why LLMs Fail: A Failure Analysis and Partial Success Measurement for Automated Security Patch Generation
von: Al-Maamari, Amir
Veröffentlicht: (2026)
von: Al-Maamari, Amir
Veröffentlicht: (2026)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
von: Evtimov, Ivan, et al.
Veröffentlicht: (2025)
Secure On-Device Video OOD Detection Without Backpropagation
von: Li, Shawn, et al.
Veröffentlicht: (2025)
von: Li, Shawn, et al.
Veröffentlicht: (2025)
ZORRO: Zero-Knowledge Robustness and Privacy for Split Learning (Full Version)
von: Sheybani, Nojan, et al.
Veröffentlicht: (2025)
von: Sheybani, Nojan, et al.
Veröffentlicht: (2025)
Fooling SHAP with Output Shuffling Attacks
von: Yuan, Jun, et al.
Veröffentlicht: (2024)
von: Yuan, Jun, et al.
Veröffentlicht: (2024)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
AI Agents May Always Fall for Prompt Injections
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2026)
Security of AI Agents
von: He, Yifeng, et al.
Veröffentlicht: (2024)
von: He, Yifeng, et al.
Veröffentlicht: (2024)
Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
von: Schmotz, David, et al.
Veröffentlicht: (2026)
von: Schmotz, David, et al.
Veröffentlicht: (2026)
Firewalls to Secure Dynamic LLM Agentic Networks
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2025)
von: Abdelnabi, Sahar, et al.
Veröffentlicht: (2025)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
von: Vyas, Sanyam, et al.
Veröffentlicht: (2024)
von: Vyas, Sanyam, et al.
Veröffentlicht: (2024)
Less is more? Rewards in RL for Cyber Defence
von: Bates, Elizabeth, et al.
Veröffentlicht: (2025)
von: Bates, Elizabeth, et al.
Veröffentlicht: (2025)
Provably Secure Agent Guardrail
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
von: Wu, Benlong, et al.
Veröffentlicht: (2026)
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Offensive Security for AI Systems: Concepts, Practices, and Applications
von: Harguess, Josh, et al.
Veröffentlicht: (2025)
von: Harguess, Josh, et al.
Veröffentlicht: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
von: Shao, Minghao, et al.
Veröffentlicht: (2025)
Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2026)
Parallax: Why AI Agents That Think Must Never Act
von: Fokou, Joel
Veröffentlicht: (2026)
von: Fokou, Joel
Veröffentlicht: (2026)
Agent Security is a Systems Problem
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2026)
von: Christodorescu, Mihai, et al.
Veröffentlicht: (2026)
Security of Internet of Agents: Attacks and Countermeasures
von: Wang, Yuntao, et al.
Veröffentlicht: (2025)
von: Wang, Yuntao, et al.
Veröffentlicht: (2025)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
von: Hong, Yining, et al.
Veröffentlicht: (2026)
von: Hong, Yining, et al.
Veröffentlicht: (2026)
aCAPTCHA: Verifying That an Entity Is a Capable Agent via Asymmetric Hardness
von: Xu, Zuyao, et al.
Veröffentlicht: (2026)
von: Xu, Zuyao, et al.
Veröffentlicht: (2026)
Large Language Models for Security Operations Centers: A Comprehensive Survey
von: Habibzadeh, Ali, et al.
Veröffentlicht: (2025)
von: Habibzadeh, Ali, et al.
Veröffentlicht: (2025)
RAG Security and Privacy: Formalizing the Threat Model and Attack Surface
von: Arzanipour, Atousa, et al.
Veröffentlicht: (2025)
von: Arzanipour, Atousa, et al.
Veröffentlicht: (2025)
Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
von: Puppala, Sai, et al.
Veröffentlicht: (2026)
von: Puppala, Sai, et al.
Veröffentlicht: (2026)
AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents
von: Zhang, Yixiang, et al.
Veröffentlicht: (2026)
von: Zhang, Yixiang, et al.
Veröffentlicht: (2026)
Agent Audit: A Security Analysis System for LLM Agent Applications
von: Zhang, Haiyue, et al.
Veröffentlicht: (2026)
von: Zhang, Haiyue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
von: Salem, Ahmed, et al.
Veröffentlicht: (2026) -
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
von: Emerson, Harry, et al.
Veröffentlicht: (2024) -
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
von: Chen, Qirui, et al.
Veröffentlicht: (2026) -
No More, No Less: Task Alignment in Terminal Agents
von: Mavali, Sina, et al.
Veröffentlicht: (2026) -
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
von: Nakamura, Mason, et al.
Veröffentlicht: (2025)