Autonomous LLM Agents & CTFs: A Second Look

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bouchari, Youness, Boffa, Matteo, Mellia, Marco, Drago, Idilio, Bui, Thanh Minh, Rossi, Dario
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911703179984896
author Bouchari, Youness
Boffa, Matteo
Mellia, Marco
Drago, Idilio
Bui, Thanh Minh
Rossi, Dario
author_facet Bouchari, Youness
Boffa, Matteo
Mellia, Marco
Drago, Idilio
Bui, Thanh Minh
Rossi, Dario
contents Large Language Model (LLM) agents are increasingly proposed to automate offensive security tasks, with recent studies reporting near human-level success rates in Capture-the-Flag (CTF) challenges. We here revisit these results, providing a second look at these claims. We engineer different agent architectures of increasing complexity and modularity on 30 web-based CTFs challenges spanning 14 vulnerability classes. We instantiate these agents with multiple LLM backbones, and compare them with claude-code, a general-purpose agent that automatically determines its internal architecture. Our evaluation yields three main findings. First, claude-code achieves performance comparable to the engineered architectures (19/30 solved tasks), suggesting that general-purpose agents are strong baselines for offensive security tasks. Second, both our architectures and claude-code struggle in the same challenge categories, revealing persistent barriers that keep current agents below human-level capability. Third, by leveraging our manually designed architectures we can systematically measure the impact of additional components, finding that structured orchestration of specialized roles outperforms monolithic designs, improving run-to-run consistency, and reducing execution costs.
format Preprint
id arxiv_https___arxiv_org_abs_2605_21497
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Autonomous LLM Agents & CTFs: A Second Look
Bouchari, Youness
Boffa, Matteo
Mellia, Marco
Drago, Idilio
Bui, Thanh Minh
Rossi, Dario
Cryptography and Security
Artificial Intelligence
Large Language Model (LLM) agents are increasingly proposed to automate offensive security tasks, with recent studies reporting near human-level success rates in Capture-the-Flag (CTF) challenges. We here revisit these results, providing a second look at these claims. We engineer different agent architectures of increasing complexity and modularity on 30 web-based CTFs challenges spanning 14 vulnerability classes. We instantiate these agents with multiple LLM backbones, and compare them with claude-code, a general-purpose agent that automatically determines its internal architecture. Our evaluation yields three main findings. First, claude-code achieves performance comparable to the engineered architectures (19/30 solved tasks), suggesting that general-purpose agents are strong baselines for offensive security tasks. Second, both our architectures and claude-code struggle in the same challenge categories, revealing persistent barriers that keep current agents below human-level capability. Third, by leveraging our manually designed architectures we can systematically measure the impact of additional components, finding that structured orchestration of specialized roles outperforms monolithic designs, improving run-to-run consistency, and reducing execution costs.
title Autonomous LLM Agents & CTFs: A Second Look
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2605.21497