HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
Fuente:
arXiv
Saved in:
| Main Authors: | Ristea, Dan, Mavroudis, Vasilios |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Referential Security as a New Paradigm for AI Evaluations
by: Ristea, Dan, et al.
Published: (2026)
by: Ristea, Dan, et al.
Published: (2026)
Less is more? Rewards in RL for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2025)
by: Bates, Elizabeth, et al.
Published: (2025)
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024)
by: Emerson, Harry, et al.
Published: (2024)
An Attentive Graph Agent for Topology-Adaptive Cyber Defence
by: Sandoval, Ilya Orson, et al.
Published: (2025)
by: Sandoval, Ilya Orson, et al.
Published: (2025)
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
Contextualized AI for Cyber Defense: An Automated Survey using LLMs
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
Operationalising Cyber Risk Management Using AI: Connecting Cyber Incidents to MITRE ATT&CK Techniques, Security Controls, and Metrics
by: Sherif, Emad, et al.
Published: (2026)
by: Sherif, Emad, et al.
Published: (2026)
AI-Driven Cyber Threat Intelligence Automation
by: Shah, Shrit, et al.
Published: (2024)
by: Shah, Shrit, et al.
Published: (2024)
Mitigating Deep Reinforcement Learning Backdoors in the Neural Activation Space
by: Vyas, Sanyam, et al.
Published: (2024)
by: Vyas, Sanyam, et al.
Published: (2024)
Entity-based Reinforcement Learning for Autonomous Cyber Defence
by: Thompson, Isaac Symes, et al.
Published: (2024)
by: Thompson, Isaac Symes, et al.
Published: (2024)
AI-Driven Security in Cloud Computing: Enhancing Threat Detection, Automated Response, and Cyber Resilience
by: Shaffi, Shamnad Mohamed, et al.
Published: (2025)
by: Shaffi, Shamnad Mohamed, et al.
Published: (2025)
ARCeR: an Agentic RAG for the Automated Definition of Cyber Ranges
by: Lupinacci, Matteo, et al.
Published: (2025)
by: Lupinacci, Matteo, et al.
Published: (2025)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment
by: Tang, Luoxi, et al.
Published: (2025)
by: Tang, Luoxi, et al.
Published: (2025)
Analysis of Publicly Accessible Operational Technology and Associated Risks
by: Rodda, Matthew, et al.
Published: (2025)
by: Rodda, Matthew, et al.
Published: (2025)
Autonomous Network Defence using Reinforcement Learning
by: Foley, Myles, et al.
Published: (2024)
by: Foley, Myles, et al.
Published: (2024)
CyberSentinel: An Emergent Threat Detection System for AI Security
by: Tallam, Krti
Published: (2025)
by: Tallam, Krti
Published: (2025)
The Path To Autonomous Cyber Defense
by: Oesch, Sean, et al.
Published: (2024)
by: Oesch, Sean, et al.
Published: (2024)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
Accelerating AI Development with Cyber Arenas
by: Cashman, William, et al.
Published: (2025)
by: Cashman, William, et al.
Published: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
by: Campbell, David, et al.
Published: (2026)
by: Campbell, David, et al.
Published: (2026)
The Security Cost of Intelligence: AI Capability, Cyber Risk, and Deployment Paradox
by: Choi, Sukwoong
Published: (2026)
by: Choi, Sukwoong
Published: (2026)
Data-Driven Falsification of Cyber-Physical Systems
by: Kundu, Atanu, et al.
Published: (2025)
by: Kundu, Atanu, et al.
Published: (2025)
Building Better Environments for Autonomous Cyber Defence
by: Hicks, Chris, et al.
Published: (2026)
by: Hicks, Chris, et al.
Published: (2026)
Cyber Threat Intelligence for Artificial Intelligence Systems
by: Krawczyk, Natalia, et al.
Published: (2026)
by: Krawczyk, Natalia, et al.
Published: (2026)
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
by: Vyas, Sanyam, et al.
Published: (2025)
by: Vyas, Sanyam, et al.
Published: (2025)
Towards Explainable and Lightweight AI for Real-Time Cyber Threat Hunting in Edge Networks
by: Rahmati, Milad
Published: (2025)
by: Rahmati, Milad
Published: (2025)
CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning
by: ElZemity, Adel, et al.
Published: (2025)
by: ElZemity, Adel, et al.
Published: (2025)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
From Incomplete Architecture to Quantified Risk: Multimodal LLM-Driven Security Assessment for Cyber-Physical Systems
by: Huang, Shaofei, et al.
Published: (2026)
by: Huang, Shaofei, et al.
Published: (2026)
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
by: Carauleanu, Marc, et al.
Published: (2024)
by: Carauleanu, Marc, et al.
Published: (2024)
Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
Similar Items
-
Referential Security as a New Paradigm for AI Evaluations
by: Ristea, Dan, et al.
Published: (2026) -
Less is more? Rewards in RL for Cyber Defence
by: Bates, Elizabeth, et al.
Published: (2025) -
CybORG++: An Enhanced Gym for the Development of Autonomous Cyber Agents
by: Emerson, Harry, et al.
Published: (2024) -
An Attentive Graph Agent for Topology-Adaptive Cyber Defence
by: Sandoval, Ilya Orson, et al.
Published: (2025) -
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)