PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Zicheng, Huang, Lige, Zhang, Jie, Liu, Dongrui, Tian, Yuan, Shao, Jing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911207744602112
author Liu, Zicheng
Huang, Lige
Zhang, Jie
Liu, Dongrui
Tian, Yuan
Shao, Jing
author_facet Liu, Zicheng
Huang, Lige
Zhang, Jie
Liu, Dongrui
Tian, Yuan
Shao, Jing
contents The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PACEbench, a practical AI cyber-exploitation benchmark built on the principles of realistic vulnerability difficulty, environmental complexity, and cyber defenses. Specifically, PACEbench comprises four scenarios spanning single, blended, chained, and defense vulnerability exploitations. To handle these complex challenges, we propose PACEagent, a novel agent that emulates human penetration testers by supporting multi-phase reconnaissance, analysis, and exploitation. Extensive experiments with seven frontier LLMs demonstrate that current models struggle with complex cyber scenarios, and none can bypass defenses. These findings suggest that current models do not yet pose a generalized cyber offense threat. Nonetheless, our work provides a robust benchmark to guide the trustworthy development of future models.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11688
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
Liu, Zicheng
Huang, Lige
Zhang, Jie
Liu, Dongrui
Tian, Yuan
Shao, Jing
Cryptography and Security
Artificial Intelligence
The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PACEbench, a practical AI cyber-exploitation benchmark built on the principles of realistic vulnerability difficulty, environmental complexity, and cyber defenses. Specifically, PACEbench comprises four scenarios spanning single, blended, chained, and defense vulnerability exploitations. To handle these complex challenges, we propose PACEagent, a novel agent that emulates human penetration testers by supporting multi-phase reconnaissance, analysis, and exploitation. Extensive experiments with seven frontier LLMs demonstrate that current models struggle with complex cyber scenarios, and none can bypass defenses. These findings suggest that current models do not yet pose a generalized cyber offense threat. Nonetheless, our work provides a robust benchmark to guide the trustworthy development of future models.
title PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.11688