CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rani, Nanda, Milner, Kimberly, Shao, Minghao, Udeshi, Meet, Xi, Haoran, Putrevu, Venkata Sai Charan, Aggarwal, Saksham, Shukla, Sandeep K., Krishnamurthy, Prashanth, Khorrami, Farshad, Shafique, Muhammad, Karri, Ramesh
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911701932179456
author Rani, Nanda
Milner, Kimberly
Shao, Minghao
Udeshi, Meet
Xi, Haoran
Putrevu, Venkata Sai Charan
Aggarwal, Saksham
Shukla, Sandeep K.
Krishnamurthy, Prashanth
Khorrami, Farshad
Shafique, Muhammad
Karri, Ramesh
author_facet Rani, Nanda
Milner, Kimberly
Shao, Minghao
Udeshi, Meet
Xi, Haoran
Putrevu, Venkata Sai Charan
Aggarwal, Saksham
Shukla, Sandeep K.
Krishnamurthy, Prashanth
Khorrami, Farshad
Shafique, Muhammad
Karri, Ramesh
contents Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effectively, but miss how real Capture-the-Flag (CTF) participants triage unknown surfaces, prioritize targets, and allocate effort under uncertainty. Current evaluations therefore fail to assess strategic reasoning beyond exploitation alone. To address this, we introduce \textit{CTFExplorer}, a benchmark suite that shifts offensive security evaluation toward a multi-target setting, which tests how agents explore, prioritize, and chain attacks. CTFExplorer deploys 40 web-based vulnerable services within a single environment, where agents must autonomously discover, distinguish, and exploit targets without predefined guidance. We also present a reactive multi-agent setup as a reference agent framework and develop an agent-agnostic evaluation framework that records structured reasoning traces for fine-grained assessment. This enables behavioral evaluation beyond binary flag capture, such as how agents manage target selection, handle failed hypotheses, coordinate across multiple stages, and extract security intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08023
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Rani, Nanda
Milner, Kimberly
Shao, Minghao
Udeshi, Meet
Xi, Haoran
Putrevu, Venkata Sai Charan
Aggarwal, Saksham
Shukla, Sandeep K.
Krishnamurthy, Prashanth
Khorrami, Farshad
Shafique, Muhammad
Karri, Ramesh
Cryptography and Security
Artificial Intelligence
Multiagent Systems
Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effectively, but miss how real Capture-the-Flag (CTF) participants triage unknown surfaces, prioritize targets, and allocate effort under uncertainty. Current evaluations therefore fail to assess strategic reasoning beyond exploitation alone. To address this, we introduce \textit{CTFExplorer}, a benchmark suite that shifts offensive security evaluation toward a multi-target setting, which tests how agents explore, prioritize, and chain attacks. CTFExplorer deploys 40 web-based vulnerable services within a single environment, where agents must autonomously discover, distinguish, and exploit targets without predefined guidance. We also present a reactive multi-agent setup as a reference agent framework and develop an agent-agnostic evaluation framework that records structured reasoning traces for fine-grained assessment. This enables behavioral evaluation beyond binary flag capture, such as how agents manage target selection, handle failed hypotheses, coordinate across multiple stages, and extract security intelligence.
title CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
topic Cryptography and Security
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2602.08023