Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Sanz-Gómez, María, Mayoral-Vilches, Víctor, Balassone, Francesco, Navarrete-Lozano, Luis Javier, Chavez, Cristóbal R. J. Veas, de Torres, Maite del Mundo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cybersecurity AI: The World's Top AI Agent for Security Capture-the-Flag (CTF)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
Cybersecurity AI in OT: Insights from an AI Top-10 Ranker in the Dragos OT CTF 2025
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
CAI Fluency: A Framework for Cybersecurity AI Fluency
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs
by: Balassone, Francesco, et al.
Published: (2025)
by: Balassone, Francesco, et al.
Published: (2025)
Cybersecurity AI: Hacking Consumer Robots in the AI Era
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
Cybersecurity AI: The Dangerous Gap Between Automation and Autonomy
by: Mayoral-Vilches, Víctor
Published: (2025)
by: Mayoral-Vilches, Víctor
Published: (2025)
Cybersecurity AI: Hacking the AI Hackers via Prompt Injection
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
The Cybersecurity of a Humanoid Robot
by: Mayoral-Vilches, Víctor
Published: (2025)
by: Mayoral-Vilches, Víctor
Published: (2025)
Offensive Robot Cybersecurity
by: Mayoral-Vilches, Víctor
Published: (2025)
by: Mayoral-Vilches, Víctor
Published: (2025)
Cybersecurity AI: Humanoid Robots as Attack Vectors
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
CAI: An Open, Bug Bounty-Ready Cybersecurity AI
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025)
Towards Cybersecurity SuperIntelligence (CSI): What's the best harness for cybersecurity?
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
Dynamic Cyber Ranges
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
by: Mayoral-Vilches, Víctor, et al.
Published: (2026)
Generative AI in Cybersecurity
by: Metta, Shivani, et al.
Published: (2024)
by: Metta, Shivani, et al.
Published: (2024)
Review of Generative AI Methods in Cybersecurity
by: Yigit, Yagmur, et al.
Published: (2024)
by: Yigit, Yagmur, et al.
Published: (2024)
Integrative Approaches in Cybersecurity and AI
by: Omar, Marwan
Published: (2024)
by: Omar, Marwan
Published: (2024)
Understanding Human-AI Collaboration in Cybersecurity Competitions
by: Tang, Tingxuan, et al.
Published: (2026)
by: Tang, Tingxuan, et al.
Published: (2026)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
Agentic AI for Cybersecurity: A Meta-Cognitive Architecture for Governable Autonomy
by: Kojukhov, Andrei, et al.
Published: (2026)
by: Kojukhov, Andrei, et al.
Published: (2026)
AgentCyTE: Leveraging Agentic AI to Generate Cybersecurity Training & Experimentation Scenarios
by: Rodriguez, Ana M., et al.
Published: (2025)
by: Rodriguez, Ana M., et al.
Published: (2025)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
Bridging Expertise Gaps: The Role of LLMs in Human-AI Collaboration for Cybersecurity
by: Tariq, Shahroz, et al.
Published: (2025)
by: Tariq, Shahroz, et al.
Published: (2025)
Constraint Migration: A Formal Theory of Throughput in AI Cybersecurity Pipelines
by: Phetmanee, Surasak
Published: (2026)
by: Phetmanee, Surasak
Published: (2026)
The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines
by: Vinay, Vaishali
Published: (2025)
by: Vinay, Vaishali
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
SECURE: Benchmarking Large Language Models for Cybersecurity
by: Bhusal, Dipkamal, et al.
Published: (2024)
by: Bhusal, Dipkamal, et al.
Published: (2024)
Automated Cybersecurity Compliance and Threat Response Using AI, Blockchain & Smart Contracts
by: Alevizos, Lampis, et al.
Published: (2024)
by: Alevizos, Lampis, et al.
Published: (2024)
Space Cybersecurity Norms
by: Sharfman, Peter, et al.
Published: (2023)
by: Sharfman, Peter, et al.
Published: (2023)
What is Cybersecurity in Space?
by: Mattar, Charbel, et al.
Published: (2025)
by: Mattar, Charbel, et al.
Published: (2025)
Cybersecurity as a Service
by: Morris, John, et al.
Published: (2024)
by: Morris, John, et al.
Published: (2024)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
by: Hakim, Safayat Bin, et al.
Published: (2025)
by: Hakim, Safayat Bin, et al.
Published: (2025)
AI/ML for 5G and Beyond Cybersecurity
by: Pirbhulal, Sandeep, et al.
Published: (2025)
by: Pirbhulal, Sandeep, et al.
Published: (2025)
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity
by: Jing, Pengfei, et al.
Published: (2024)
by: Jing, Pengfei, et al.
Published: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Hyperloop: A Cybersecurity Perspective
by: Brighente, Alessandro, et al.
Published: (2022)
by: Brighente, Alessandro, et al.
Published: (2022)
Cybersecurity: Past, Present and Future
by: Alam, Shahid
Published: (2022)
by: Alam, Shahid
Published: (2022)
Does Johnny Get the Message? Evaluating Cybersecurity Notifications for Everyday Users
by: Jüttner, Victor, et al.
Published: (2025)
by: Jüttner, Victor, et al.
Published: (2025)
Similar Items
-
Cybersecurity AI: The World's Top AI Agent for Security Capture-the-Flag (CTF)
by: Mayoral-Vilches, Víctor, et al.
Published: (2025) -
Cybersecurity AI in OT: Insights from an AI Top-10 Ranker in the Dragos OT CTF 2025
by: Mayoral-Vilches, Víctor, et al.
Published: (2025) -
Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense
by: Mayoral-Vilches, Víctor, et al.
Published: (2026) -
CAI Fluency: A Framework for Cybersecurity AI Fluency
by: Mayoral-Vilches, Víctor, et al.
Published: (2025) -
Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs
by: Balassone, Francesco, et al.
Published: (2025)