Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Merves, Tyler H., Conaway, Michael H., Escobar, Joseph M., Otal, Hakan T., Tatar, Unal
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917419981733888
author Merves, Tyler H.
Conaway, Michael H.
Escobar, Joseph M.
Otal, Hakan T.
Tatar, Unal
author_facet Merves, Tyler H.
Conaway, Michael H.
Escobar, Joseph M.
Otal, Hakan T.
Tatar, Unal
contents We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier models from 7 providers on all 200 challenges of the NYU CTF Bench. Building on the D-CIPHER multi-agent framework, we extend it with multi-provider backend support, a custom Kali Linux environment with over 100 pre-installed penetration testing tools, and runtime tool-discovery agents. Through a controlled factorial study, we find that the Kali Linux environment yields a +9.5 percentage-point improvement over Ubuntu, while auto-prompting and category-specific tips often degrade performance in well-equipped environments. Among models, Claude 4.5 Opus achieves the highest solve rate (59%), followed by Gemini 3 Pro (52%), with Gemini 3 Flash offering the best cost-efficiency at $0.05 per solve. Asymmetric planner/executor model assignments provide no meaningful benefit while coherent same-model configurations consistently outperform mixed-tier pairings. Our results indicate that environment tooling and model selection emerge as the strongest drivers of performance, whereas prompt engineering interventions show diminishing or negative returns in well-equipped environments. Reported performance reflects both model reasoning ability and compatibility with agent tooling and API integration.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17159
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
Merves, Tyler H.
Conaway, Michael H.
Escobar, Joseph M.
Otal, Hakan T.
Tatar, Unal
Cryptography and Security
Artificial Intelligence
Computation and Language
68M25, 68T50
K.6.5; I.2.11; I.2.7
We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier models from 7 providers on all 200 challenges of the NYU CTF Bench. Building on the D-CIPHER multi-agent framework, we extend it with multi-provider backend support, a custom Kali Linux environment with over 100 pre-installed penetration testing tools, and runtime tool-discovery agents. Through a controlled factorial study, we find that the Kali Linux environment yields a +9.5 percentage-point improvement over Ubuntu, while auto-prompting and category-specific tips often degrade performance in well-equipped environments. Among models, Claude 4.5 Opus achieves the highest solve rate (59%), followed by Gemini 3 Pro (52%), with Gemini 3 Flash offering the best cost-efficiency at $0.05 per solve. Asymmetric planner/executor model assignments provide no meaningful benefit while coherent same-model configurations consistently outperform mixed-tier pairings. Our results indicate that environment tooling and model selection emerge as the strongest drivers of performance, whereas prompt engineering interventions show diminishing or negative returns in well-equipped environments. Reported performance reflects both model reasoning ability and compatibility with agent tooling and API integration.
title Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks
topic Cryptography and Security
Artificial Intelligence
Computation and Language
68M25, 68T50
K.6.5; I.2.11; I.2.7
url https://arxiv.org/abs/2604.17159