Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Minghao, Rani, Nanda, Milner, Kimberly, Xi, Haoran, Udeshi, Meet, Aggarwal, Saksham, Putrevu, Venkata Sai Charan, Shukla, Sandeep Kumar, Krishnamurthy, Prashanth, Khorrami, Farshad, Karri, Ramesh, Shafique, Muhammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915925132836864
author Shao, Minghao
Rani, Nanda
Milner, Kimberly
Xi, Haoran
Udeshi, Meet
Aggarwal, Saksham
Putrevu, Venkata Sai Charan
Shukla, Sandeep Kumar
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
author_facet Shao, Minghao
Rani, Nanda
Milner, Kimberly
Xi, Haoran
Udeshi, Meet
Aggarwal, Saksham
Putrevu, Venkata Sai Charan
Shukla, Sandeep Kumar
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
contents Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate the key factors that drive agent success and provide a detailed recipe for building effective LLM-based offensive security agents. First, we present CTFJudge, a framework leveraging LLM as a judge to analyze agent trajectories and provide granular evaluation across CTF solving steps. Second, we propose a novel metric, CTF Competency Index (CCI) for partial correctness, revealing how closely agent solutions align with human-crafted gold standards. Third, we examine how LLM hyperparameters, namely temperature, top-p, and maximum token length, influence agent performance and automated cybersecurity task planning. For rapid evaluation, we present CTFTiny, a curated benchmark of 50 representative CTF challenges across binary exploitation, web, reverse engineering, forensics, and cryptography. Our findings identify optimal multi-agent coordination settings and lay the groundwork for future LLM agent research in cybersecurity. We make CTFTiny open source to public https://github.com/NYU-LLM-CTF/CTFTiny along with CTFJudge on https://github.com/NYU-LLM-CTF/CTFJudge.
format Preprint
id arxiv_https___arxiv_org_abs_2508_05674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
Shao, Minghao
Rani, Nanda
Milner, Kimberly
Xi, Haoran
Udeshi, Meet
Aggarwal, Saksham
Putrevu, Venkata Sai Charan
Shukla, Sandeep Kumar
Krishnamurthy, Prashanth
Khorrami, Farshad
Karri, Ramesh
Shafique, Muhammad
Cryptography and Security
Artificial Intelligence
Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate the key factors that drive agent success and provide a detailed recipe for building effective LLM-based offensive security agents. First, we present CTFJudge, a framework leveraging LLM as a judge to analyze agent trajectories and provide granular evaluation across CTF solving steps. Second, we propose a novel metric, CTF Competency Index (CCI) for partial correctness, revealing how closely agent solutions align with human-crafted gold standards. Third, we examine how LLM hyperparameters, namely temperature, top-p, and maximum token length, influence agent performance and automated cybersecurity task planning. For rapid evaluation, we present CTFTiny, a curated benchmark of 50 representative CTF challenges across binary exploitation, web, reverse engineering, forensics, and cryptography. Our findings identify optimal multi-agent coordination settings and lay the groundwork for future LLM agent research in cybersecurity. We make CTFTiny open source to public https://github.com/NYU-LLM-CTF/CTFTiny along with CTFJudge on https://github.com/NYU-LLM-CTF/CTFJudge.
title Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2508.05674