ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Seunghyun, Brumley, David |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024)
by: Jing, Pengfei, et al.
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)
by: Ma, Haokai, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
LLM Agents can Autonomously Exploit One-day Vulnerabilities
by: Fang, Richard, et al.
Published: (2024)
by: Fang, Richard, et al.
Published: (2024)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
by: Najt, Elle, et al.
Published: (2026)
by: Najt, Elle, et al.
Published: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Dynamic Risk Assessments for Offensive Cybersecurity Agents
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026)
by: Probst, Benjamin, et al.
Published: (2026)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization
by: Meidan, Yair, et al.
Published: (2026)
by: Meidan, Yair, et al.
Published: (2026)
CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents
by: Liu, Zesen, et al.
Published: (2025)
by: Liu, Zesen, et al.
Published: (2025)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
by: Ferrag, Mohamed Amine, et al.
Published: (2024)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026)
by: Fan, Yihe, et al.
Published: (2026)
Enforcing Cybersecurity Constraints for LLM-driven Robot Agents for Online Transactions
by: Shah, Shraddha Pradipbhai, et al.
Published: (2025)
by: Shah, Shraddha Pradipbhai, et al.
Published: (2025)
ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
by: Lau, Nancy, et al.
Published: (2026)
by: Lau, Nancy, et al.
Published: (2026)
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
by: Huang, Enhao, et al.
Published: (2025)
by: Huang, Enhao, et al.
Published: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
Hallucination as Exploit: Evidence-Carrying Multimodal Agents
by: Zhang, Guijia, et al.
Published: (2026)
by: Zhang, Guijia, et al.
Published: (2026)
AI Agent Smart Contract Exploit Generation
by: Gervais, Arthur, et al.
Published: (2025)
by: Gervais, Arthur, et al.
Published: (2025)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
From Legacy to Standard: LLM-Assisted Transformation of Cybersecurity Playbooks into CACAO Format
by: Gurabi, Mehdi Akbari, et al.
Published: (2025)
by: Gurabi, Mehdi Akbari, et al.
Published: (2025)
Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents
by: Jiang, Linxi, et al.
Published: (2026)
by: Jiang, Linxi, et al.
Published: (2026)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
by: Zou, Wei, et al.
Published: (2026)
by: Zou, Wei, et al.
Published: (2026)
RedSage: A Cybersecurity Generalist LLM
by: Suryanto, Naufal, et al.
Published: (2026)
by: Suryanto, Naufal, et al.
Published: (2026)
Poster: SpiderSim: Multi-Agent Driven Theoretical Cybersecurity Simulation for Industrial Digitalization
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
DPImageBench: A Unified Benchmark for Differentially Private Image Synthesis
by: Gong, Chen, et al.
Published: (2025)
by: Gong, Chen, et al.
Published: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
by: Dahiya, Vivek, et al.
Published: (2026)
by: Dahiya, Vivek, et al.
Published: (2026)
Exploiting LLM Quantization
by: Egashira, Kazuki, et al.
Published: (2024)
by: Egashira, Kazuki, et al.
Published: (2024)
Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
by: Lian, Zhuotao, et al.
Published: (2025)
by: Lian, Zhuotao, et al.
Published: (2025)
Towards the Development of an LLM-Based Methodology for Automated Security Profiling in Compliance with Ukrainian Cybersecurity Regulations
by: Shafranskyi, Daniil, et al.
Published: (2026)
by: Shafranskyi, Daniil, et al.
Published: (2026)
Similar Items
-
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024) -
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024) -
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026) -
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024) -
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
by: Ma, Haokai, et al.
Published: (2025)