From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Lanxiao, Dave, Daksh, Cody, Tyler, Beling, Peter, Jin, Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911262077616128
author Huang, Lanxiao
Dave, Daksh
Cody, Tyler
Beling, Peter
Jin, Ming
author_facet Huang, Lanxiao
Dave, Daksh
Cody, Tyler
Beling, Peter
Jin, Ming
contents Large language models (LLMs) are increasingly used to automate or augment penetration testing, but their effectiveness and reliability across attack phases remain unclear. We present a comprehensive evaluation of multiple LLM-based agents, from single-agent to modular designs, across realistic penetration testing scenarios, measuring empirical performance and recurring failure patterns. We also isolate the impact of five core functional capabilities via targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM). These interventions support, respectively: (i) context coherence and retention, (ii) inter-component coordination and state management, (iii) tool use accuracy and selective execution, (iv) multi-step strategic planning, error detection, and recovery, and (v) real-time dynamic responsiveness. Our results show that while some architectures natively exhibit subsets of these properties, targeted augmentations substantially improve modular agent performance, especially in complex, multi-step, and real-time penetration testing tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14289
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
Huang, Lanxiao
Dave, Daksh
Cody, Tyler
Beling, Peter
Jin, Ming
Artificial Intelligence
Computation and Language
Machine Learning
68T50
I.2.7; I.5.4
Large language models (LLMs) are increasingly used to automate or augment penetration testing, but their effectiveness and reliability across attack phases remain unclear. We present a comprehensive evaluation of multiple LLM-based agents, from single-agent to modular designs, across realistic penetration testing scenarios, measuring empirical performance and recurring failure patterns. We also isolate the impact of five core functional capabilities via targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM). These interventions support, respectively: (i) context coherence and retention, (ii) inter-component coordination and state management, (iii) tool use accuracy and selective execution, (iv) multi-step strategic planning, error detection, and recovery, and (v) real-time dynamic responsiveness. Our results show that while some architectures natively exhibit subsets of these properties, targeted augmentations substantially improve modular agent performance, especially in complex, multi-step, and real-time penetration testing tasks.
title From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
topic Artificial Intelligence
Computation and Language
Machine Learning
68T50
I.2.7; I.5.4
url https://arxiv.org/abs/2509.14289