Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Anurin, Andrey, Ng, Jonathan, Schaffer, Kibo, Schreiber, Jason, Kran, Esben |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Uplifted Attackers, Human Defenders: The Cyber Offense-Defense Balance for Trailing-Edge Organizations
by: Murphy, Benjamin, et al.
Published: (2025)
by: Murphy, Benjamin, et al.
Published: (2025)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing
by: Hu, Jinwei, et al.
Published: (2026)
by: Hu, Jinwei, et al.
Published: (2026)
When Scanners Lie: Evaluator Instability in LLM Red-Teaming
by: Erez, Lidor, et al.
Published: (2026)
by: Erez, Lidor, et al.
Published: (2026)
Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand
by: Koch, Christopher
Published: (2026)
by: Koch, Christopher
Published: (2026)
Cyber Deception Reactive: TCP Stealth Redirection to On-Demand Honeypots
by: Lopez, Pedro Beltran, et al.
Published: (2024)
by: Lopez, Pedro Beltran, et al.
Published: (2024)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
by: Lee, Seunghyun, et al.
Published: (2026)
by: Lee, Seunghyun, et al.
Published: (2026)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
The Security Cost of Intelligence: AI Capability, Cyber Risk, and Deployment Paradox
by: Choi, Sukwoong
Published: (2026)
by: Choi, Sukwoong
Published: (2026)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
by: Afane, Khalifa, et al.
Published: (2024)
by: Afane, Khalifa, et al.
Published: (2024)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
Breaking the Loop: Detecting and Mitigating Denial-of-Service Vulnerabilities in Large Language Models
by: Yu, Junzhe, et al.
Published: (2025)
by: Yu, Junzhe, et al.
Published: (2025)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
by: Bhardwaj, Arth, et al.
Published: (2025)
by: Bhardwaj, Arth, et al.
Published: (2025)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
Evaluation of Hash Algorithm Performance for Cryptocurrency Exchanges Based on Blockchain System
by: Chen, Abel C. H.
Published: (2024)
by: Chen, Abel C. H.
Published: (2024)
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
by: Meng, Yuqiao, et al.
Published: (2025)
by: Meng, Yuqiao, et al.
Published: (2025)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
by: Marchand, Rahul, et al.
Published: (2026)
by: Marchand, Rahul, et al.
Published: (2026)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
by: Huang, Enhao, et al.
Published: (2025)
by: Huang, Enhao, et al.
Published: (2025)
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026)
by: Probst, Benjamin, et al.
Published: (2026)
Quantifying Loss Aversion in Cyber Adversaries via LLM Analysis
by: Hans, Soham, et al.
Published: (2025)
by: Hans, Soham, et al.
Published: (2025)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
by: Fan, Yihe, et al.
Published: (2026)
by: Fan, Yihe, et al.
Published: (2026)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
by: Ayzenshteyn, Daniel, et al.
Published: (2024)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
by: Black, Sid, et al.
Published: (2025)
by: Black, Sid, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
The Path To Autonomous Cyber Defense
by: Oesch, Sean, et al.
Published: (2024)
by: Oesch, Sean, et al.
Published: (2024)
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
BARTPredict: Empowering IoT Security with LLM-Driven Cyber Threat Prediction
by: Diaf, Alaeddine, et al.
Published: (2025)
by: Diaf, Alaeddine, et al.
Published: (2025)
ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation
by: Wu, Yiran, et al.
Published: (2025)
by: Wu, Yiran, et al.
Published: (2025)
Similar Items
-
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
by: Li, Yu, et al.
Published: (2025) -
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025) -
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025) -
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025) -
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)