PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zicheng, Huang, Lige, Zhang, Jie, Liu, Dongrui, Tian, Yuan, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025)
by: Ji, Zimo, et al.
Published: (2025)
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025)
by: Rodriguez, Mikel, et al.
Published: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Emerging Cyber Attack Risks of Medical AI Agents
by: Qiu, Jianing, et al.
Published: (2025)
by: Qiu, Jianing, et al.
Published: (2025)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
by: Deason, Lauren, et al.
Published: (2025)
by: Deason, Lauren, et al.
Published: (2025)
A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
by: Wang, Jinghao, et al.
Published: (2025)
by: Wang, Jinghao, et al.
Published: (2025)
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
by: Yuan, Leitao, et al.
Published: (2026)
by: Yuan, Leitao, et al.
Published: (2026)
IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation
by: Guo, Yanpei, et al.
Published: (2026)
by: Guo, Yanpei, et al.
Published: (2026)
The Security Cost of Intelligence: AI Capability, Cyber Risk, and Deployment Paradox
by: Choi, Sukwoong
Published: (2026)
by: Choi, Sukwoong
Published: (2026)
A Trajectory-Based Safety Audit of Clawdbot (OpenClaw)
by: Chen, Tianyu, et al.
Published: (2026)
by: Chen, Tianyu, et al.
Published: (2026)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
CyberForce: A Federated Reinforcement Learning Framework for Malware Mitigation
by: Feng, Chao, et al.
Published: (2023)
by: Feng, Chao, et al.
Published: (2023)
CyberSentinel: An Emergent Threat Detection System for AI Security
by: Tallam, Krti
Published: (2025)
by: Tallam, Krti
Published: (2025)
Contextualized AI for Cyber Defense: An Automated Survey using LLMs
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
by: Haryanto, Christoforus Yoga, et al.
Published: (2024)
Operationalising Cyber Risk Management Using AI: Connecting Cyber Incidents to MITRE ATT&CK Techniques, Security Controls, and Metrics
by: Sherif, Emad, et al.
Published: (2026)
by: Sherif, Emad, et al.
Published: (2026)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
by: Bhardwaj, Arth, et al.
Published: (2025)
by: Bhardwaj, Arth, et al.
Published: (2025)
CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge
by: Keppler, Gustav, et al.
Published: (2026)
by: Keppler, Gustav, et al.
Published: (2026)
AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
Quantifying Association Capabilities of Large Language Models and Its Implications on Privacy Leakage
by: Shao, Hanyin, et al.
Published: (2023)
by: Shao, Hanyin, et al.
Published: (2023)
AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises
by: Tung, Ivan K., et al.
Published: (2026)
by: Tung, Ivan K., et al.
Published: (2026)
AI-Augmented Ethical Hacking: A Practical Examination of Manual Exploitation and Privilege Escalation in Linux Environments
by: Al-Sinani, Haitham S., et al.
Published: (2024)
by: Al-Sinani, Haitham S., et al.
Published: (2024)
SoK: DARPA's AI Cyber Challenge (AIxCC): Competition Design, Architectures, and Lessons Learned
by: Zhang, Cen, et al.
Published: (2026)
by: Zhang, Cen, et al.
Published: (2026)
Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
A Large Language Model-Supported Threat Modeling Framework for Transportation Cyber-Physical Systems
by: Salek, M Sabbir, et al.
Published: (2025)
by: Salek, M Sabbir, et al.
Published: (2025)
Large Language Models for Cyber Security: A Systematic Literature Review
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval
by: Hamzic, Dzenan, et al.
Published: (2026)
by: Hamzic, Dzenan, et al.
Published: (2026)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Towards Explainable and Lightweight AI for Real-Time Cyber Threat Hunting in Edge Networks
by: Rahmati, Milad
Published: (2025)
by: Rahmati, Milad
Published: (2025)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
by: Lohn, Andrew J.
Published: (2025)
by: Lohn, Andrew J.
Published: (2025)
CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence
by: Cheng, Yutong, et al.
Published: (2025)
by: Cheng, Yutong, et al.
Published: (2025)
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge
by: Tihanyi, Norbert, et al.
Published: (2024)
by: Tihanyi, Norbert, et al.
Published: (2024)
EdgeShield: A Universal and Efficient Edge Computing Framework for Robust AI
by: Zhong, Duo, et al.
Published: (2024)
by: Zhong, Duo, et al.
Published: (2024)
NetMoniAI: An Agentic AI Framework for Network Security & Monitoring
by: Zambare, Pallavi, et al.
Published: (2025)
by: Zambare, Pallavi, et al.
Published: (2025)
Similar Items
-
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026) -
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024) -
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
by: Ji, Zimo, et al.
Published: (2025) -
A Framework for Evaluating Emerging Cyberattack Capabilities of AI
by: Rodriguez, Mikel, et al.
Published: (2025) -
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)