CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhun, Shi, Tianneng, He, Jingxuan, Cai, Matthew, Zhang, Jialin, Song, Dawn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026)
by: Wang, Zhun, et al.
Published: (2026)
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)
by: Siu, Vincent, et al.
Published: (2026)
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
by: Tang, Yuheng, et al.
Published: (2026)
by: Tang, Yuheng, et al.
Published: (2026)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026)
by: Fan, Yihe, et al.
Published: (2026)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
by: Kim, Juhee, et al.
Published: (2026)
by: Kim, Juhee, et al.
Published: (2026)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity
by: Roy, Shovan
Published: (2025)
by: Roy, Shovan
Published: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
PromptArmor: Simple yet Effective Prompt Injection Defenses
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
by: Zhang, Andy K., et al.
Published: (2025)
by: Zhang, Andy K., et al.
Published: (2025)
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
by: Nie, Yuzhou, et al.
Published: (2025)
by: Nie, Yuzhou, et al.
Published: (2025)
SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
by: Conde, Pedro, et al.
Published: (2026)
by: Conde, Pedro, et al.
Published: (2026)
OpenSage: Self-programming Agent Generation Engine
by: Li, Hongwei, et al.
Published: (2026)
by: Li, Hongwei, et al.
Published: (2026)
Integrative Approaches in Cybersecurity and AI
by: Omar, Marwan
Published: (2024)
by: Omar, Marwan
Published: (2024)
Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
by: Hakim, Safayat Bin, et al.
Published: (2025)
by: Hakim, Safayat Bin, et al.
Published: (2025)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
by: Bhardwaj, Arth, et al.
Published: (2025)
by: Bhardwaj, Arth, et al.
Published: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Dynamic Risk Assessments for Offensive Cybersecurity Agents
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
Towards Explainable and Lightweight AI for Real-Time Cyber Threat Hunting in Edge Networks
by: Rahmati, Milad
Published: (2025)
by: Rahmati, Milad
Published: (2025)
OSS-CRS: Liberating AIxCC Cyber Reasoning Systems for Real-World Open-Source Security
by: Chin, Andrew, et al.
Published: (2026)
by: Chin, Andrew, et al.
Published: (2026)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025)
by: Rodriguez, Mikel, et al.
Published: (2025)
Artificial Intelligence in Cybersecurity: Building Resilient Cyber Diplomacy Frameworks
by: Stoltz, Michael
Published: (2024)
by: Stoltz, Michael
Published: (2024)
AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises
by: Tung, Ivan K., et al.
Published: (2026)
by: Tung, Ivan K., et al.
Published: (2026)
Similar Items
-
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025) -
Frontier AI's Impact on the Cybersecurity Landscape
by: Potter, Yujin, et al.
Published: (2025) -
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
by: Wang, Zhun, et al.
Published: (2026) -
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
by: Wang, Zhun, et al.
Published: (2025) -
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)