SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Fuente:
arXiv
Guardado en:
| Autores principales: | Najt, Elle, Toft, Colin, Tracy, Tyler, Roger, Fabien, Benton, Joe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
por: Jotautaitė, Monika, et al.
Publicado: (2026)
por: Jotautaitė, Monika, et al.
Publicado: (2026)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
por: Kang, Mintong, et al.
Publicado: (2023)
por: Kang, Mintong, et al.
Publicado: (2023)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
por: Zhang, Dongsen, et al.
Publicado: (2025)
por: Zhang, Dongsen, et al.
Publicado: (2025)
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
por: Domico, Kyle, et al.
Publicado: (2025)
por: Domico, Kyle, et al.
Publicado: (2025)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
por: Kutasov, Jonathan, et al.
Publicado: (2025)
por: Kutasov, Jonathan, et al.
Publicado: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
por: Zhang, Hanrong, et al.
Publicado: (2024)
por: Zhang, Hanrong, et al.
Publicado: (2024)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
por: Evtimov, Ivan, et al.
Publicado: (2025)
por: Evtimov, Ivan, et al.
Publicado: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
por: Lim, Taein, et al.
Publicado: (2026)
por: Lim, Taein, et al.
Publicado: (2026)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
por: Ma, Haokai, et al.
Publicado: (2025)
por: Ma, Haokai, et al.
Publicado: (2025)
EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
por: Liu, Shigang, et al.
Publicado: (2024)
por: Liu, Shigang, et al.
Publicado: (2024)
Reliable Model Watermarking: Defending Against Theft without Compromising on Evasion
por: Zhu, Hongyu, et al.
Publicado: (2024)
por: Zhu, Hongyu, et al.
Publicado: (2024)
BashArena: A Control Setting for Highly Privileged AI Agents
por: Kaufman, Adam, et al.
Publicado: (2025)
por: Kaufman, Adam, et al.
Publicado: (2025)
Securing AI Agents Against Prompt Injection Attacks
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
por: Ramakrishnan, Badrinath, et al.
Publicado: (2025)
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
por: Schaeffer, Joachim, et al.
Publicado: (2026)
por: Schaeffer, Joachim, et al.
Publicado: (2026)
MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks
por: Zhang, Xinkai, et al.
Publicado: (2026)
por: Zhang, Xinkai, et al.
Publicado: (2026)
AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
por: Wu, Yixin, et al.
Publicado: (2025)
por: Wu, Yixin, et al.
Publicado: (2025)
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
por: Cao, Tri, et al.
Publicado: (2025)
por: Cao, Tri, et al.
Publicado: (2025)
ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language Models
por: Liu, Xuxu, et al.
Publicado: (2025)
por: Liu, Xuxu, et al.
Publicado: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
por: Lee, Seunghyun, et al.
Publicado: (2026)
por: Lee, Seunghyun, et al.
Publicado: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
por: Gioacchini, Luca, et al.
Publicado: (2024)
por: Gioacchini, Luca, et al.
Publicado: (2024)
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
por: Xiang, Qiuchi, et al.
Publicado: (2026)
por: Xiang, Qiuchi, et al.
Publicado: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
por: Wang, Zhilong, et al.
Publicado: (2025)
por: Wang, Zhilong, et al.
Publicado: (2025)
Narrow Secret Loyalty Dodges Black-Box Audits
por: Lamerton, Alfie, et al.
Publicado: (2026)
por: Lamerton, Alfie, et al.
Publicado: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
por: Zhu, Kaijie, et al.
Publicado: (2025)
por: Zhu, Kaijie, et al.
Publicado: (2025)
Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks
por: G, Aravindhan, et al.
Publicado: (2025)
por: G, Aravindhan, et al.
Publicado: (2025)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
por: Yin, Sheng, et al.
Publicado: (2024)
por: Yin, Sheng, et al.
Publicado: (2024)
Enhancing O-RAN Security: Evasion Attacks and Robust Defenses for Graph Reinforcement Learning-based Connection Management
por: Balakrishnan, Ravikumar, et al.
Publicado: (2024)
por: Balakrishnan, Ravikumar, et al.
Publicado: (2024)
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
por: Zhao, Lei, et al.
Publicado: (2026)
por: Zhao, Lei, et al.
Publicado: (2026)
Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
por: Xu, Wenpeng
Publicado: (2026)
por: Xu, Wenpeng
Publicado: (2026)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
por: Li, Siyuan, et al.
Publicado: (2026)
por: Li, Siyuan, et al.
Publicado: (2026)
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
por: Syros, Georgios, et al.
Publicado: (2026)
por: Syros, Georgios, et al.
Publicado: (2026)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
por: Yueh-Han, Chen, et al.
Publicado: (2025)
por: Yueh-Han, Chen, et al.
Publicado: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
por: Shen, Chihao, et al.
Publicado: (2025)
por: Shen, Chihao, et al.
Publicado: (2025)
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
por: Xiang, Chong, et al.
Publicado: (2026)
por: Xiang, Chong, et al.
Publicado: (2026)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
por: Liu, Shi, et al.
Publicado: (2026)
por: Liu, Shi, et al.
Publicado: (2026)
LLM Watermark Evasion via Bias Inversion
por: Hwang, Jeongyeon, et al.
Publicado: (2025)
por: Hwang, Jeongyeon, et al.
Publicado: (2025)
A Model Stealing Attack Against Multi-Exit Networks
por: Pan, Li, et al.
Publicado: (2023)
por: Pan, Li, et al.
Publicado: (2023)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
por: Jaiswal, Piyush, et al.
Publicado: (2026)
por: Jaiswal, Piyush, et al.
Publicado: (2026)
Defenses Against Prompt Attacks Learn Surface Heuristics
por: Li, Shawn, et al.
Publicado: (2026)
por: Li, Shawn, et al.
Publicado: (2026)
Ejemplares similares
-
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
por: Jotautaitė, Monika, et al.
Publicado: (2026) -
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
por: Kang, Mintong, et al.
Publicado: (2023) -
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
por: Zhang, Dongsen, et al.
Publicado: (2025) -
Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning
por: Domico, Kyle, et al.
Publicado: (2025) -
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
por: Kutasov, Jonathan, et al.
Publicado: (2025)