CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yuxuan, Kellermann, Antony, Bowman, Dylan, Li, Philip, Gupta, Akul, Danda, Adarsh, Fang, Richard, Jensen, Conner, Ihli, Eric, Benn, Jason, Geronimo, Jet, Dhir, Avi, Rao, Sudhit, Yu, Kaicheng, Stone, Twm, Kang, Daniel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909657875873792
author Zhu, Yuxuan
Kellermann, Antony
Bowman, Dylan
Li, Philip
Gupta, Akul
Danda, Adarsh
Fang, Richard
Jensen, Conner
Ihli, Eric
Benn, Jason
Geronimo, Jet
Dhir, Avi
Rao, Sudhit
Yu, Kaicheng
Stone, Twm
Kang, Daniel
author_facet Zhu, Yuxuan
Kellermann, Antony
Bowman, Dylan
Li, Philip
Gupta, Akul
Danda, Adarsh
Fang, Richard
Jensen, Conner
Ihli, Eric
Benn, Jason
Geronimo, Jet
Dhir, Avi
Rao, Sudhit
Yu, Kaicheng
Stone, Twm
Kang, Daniel
contents Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture the Flag competitions or lack comprehensive coverage. Building a benchmark for real-world vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable threats. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our evaluation shows that the state-of-the-art agent framework can resolve up to 13% of vulnerabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17332
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Zhu, Yuxuan
Kellermann, Antony
Bowman, Dylan
Li, Philip
Gupta, Akul
Danda, Adarsh
Fang, Richard
Jensen, Conner
Ihli, Eric
Benn, Jason
Geronimo, Jet
Dhir, Avi
Rao, Sudhit
Yu, Kaicheng
Stone, Twm
Kang, Daniel
Cryptography and Security
Artificial Intelligence
I.2.1; I.2.7
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the ability of LLM agents to exploit web application vulnerabilities. However, existing benchmarks fall short as they are limited to abstracted Capture the Flag competitions or lack comprehensive coverage. Building a benchmark for real-world vulnerabilities involves both specialized expertise to reproduce exploits and a systematic approach to evaluating unpredictable threats. To address this challenge, we introduce CVE-Bench, a real-world cybersecurity benchmark based on critical-severity Common Vulnerabilities and Exposures. In CVE-Bench, we design a sandbox framework that enables LLM agents to exploit vulnerable web applications in scenarios that mimic real-world conditions, while also providing effective evaluation of their exploits. Our evaluation shows that the state-of-the-art agent framework can resolve up to 13% of vulnerabilities.
title CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
topic Cryptography and Security
Artificial Intelligence
I.2.1; I.2.7
url https://arxiv.org/abs/2503.17332