CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911701932179456 |
|---|---|
| author | Rani, Nanda Milner, Kimberly Shao, Minghao Udeshi, Meet Xi, Haoran Putrevu, Venkata Sai Charan Aggarwal, Saksham Shukla, Sandeep K. Krishnamurthy, Prashanth Khorrami, Farshad Shafique, Muhammad Karri, Ramesh |
| author_facet | Rani, Nanda Milner, Kimberly Shao, Minghao Udeshi, Meet Xi, Haoran Putrevu, Venkata Sai Charan Aggarwal, Saksham Shukla, Sandeep K. Krishnamurthy, Prashanth Khorrami, Farshad Shafique, Muhammad Karri, Ramesh |
| contents | Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effectively, but miss how real Capture-the-Flag (CTF) participants triage unknown surfaces, prioritize targets, and allocate effort under uncertainty. Current evaluations therefore fail to assess strategic reasoning beyond exploitation alone. To address this, we introduce \textit{CTFExplorer}, a benchmark suite that shifts offensive security evaluation toward a multi-target setting, which tests how agents explore, prioritize, and chain attacks. CTFExplorer deploys 40 web-based vulnerable services within a single environment, where agents must autonomously discover, distinguish, and exploit targets without predefined guidance. We also present a reactive multi-agent setup as a reference agent framework and develop an agent-agnostic evaluation framework that records structured reasoning traces for fine-grained assessment. This enables behavioral evaluation beyond binary flag capture, such as how agents manage target selection, handle failed hypotheses, coordinate across multiple stages, and extract security intelligence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_08023 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking Rani, Nanda Milner, Kimberly Shao, Minghao Udeshi, Meet Xi, Haoran Putrevu, Venkata Sai Charan Aggarwal, Saksham Shukla, Sandeep K. Krishnamurthy, Prashanth Khorrami, Farshad Shafique, Muhammad Karri, Ramesh Cryptography and Security Artificial Intelligence Multiagent Systems Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effectively, but miss how real Capture-the-Flag (CTF) participants triage unknown surfaces, prioritize targets, and allocate effort under uncertainty. Current evaluations therefore fail to assess strategic reasoning beyond exploitation alone. To address this, we introduce \textit{CTFExplorer}, a benchmark suite that shifts offensive security evaluation toward a multi-target setting, which tests how agents explore, prioritize, and chain attacks. CTFExplorer deploys 40 web-based vulnerable services within a single environment, where agents must autonomously discover, distinguish, and exploit targets without predefined guidance. We also present a reactive multi-agent setup as a reference agent framework and develop an agent-agnostic evaluation framework that records structured reasoning traces for fine-grained assessment. This enables behavioral evaluation beyond binary flag capture, such as how agents manage target selection, handle failed hypotheses, coordinate across multiple stages, and extract security intelligence. |
| title | CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking |
| topic | Cryptography and Security Artificial Intelligence Multiagent Systems |
| url | https://arxiv.org/abs/2602.08023 |