From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911262077616128 |
|---|---|
| author | Huang, Lanxiao Dave, Daksh Cody, Tyler Beling, Peter Jin, Ming |
| author_facet | Huang, Lanxiao Dave, Daksh Cody, Tyler Beling, Peter Jin, Ming |
| contents | Large language models (LLMs) are increasingly used to automate or augment penetration testing, but their effectiveness and reliability across attack phases remain unclear. We present a comprehensive evaluation of multiple LLM-based agents, from single-agent to modular designs, across realistic penetration testing scenarios, measuring empirical performance and recurring failure patterns. We also isolate the impact of five core functional capabilities via targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM). These interventions support, respectively: (i) context coherence and retention, (ii) inter-component coordination and state management, (iii) tool use accuracy and selective execution, (iv) multi-step strategic planning, error detection, and recovery, and (v) real-time dynamic responsiveness. Our results show that while some architectures natively exhibit subsets of these properties, targeted augmentations substantially improve modular agent performance, especially in complex, multi-step, and real-time penetration testing tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_14289 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing Huang, Lanxiao Dave, Daksh Cody, Tyler Beling, Peter Jin, Ming Artificial Intelligence Computation and Language Machine Learning 68T50 I.2.7; I.5.4 Large language models (LLMs) are increasingly used to automate or augment penetration testing, but their effectiveness and reliability across attack phases remain unclear. We present a comprehensive evaluation of multiple LLM-based agents, from single-agent to modular designs, across realistic penetration testing scenarios, measuring empirical performance and recurring failure patterns. We also isolate the impact of five core functional capabilities via targeted augmentations: Global Context Memory (GCM), Inter-Agent Messaging (IAM), Context-Conditioned Invocation (CCI), Adaptive Planning (AP), and Real-Time Monitoring (RTM). These interventions support, respectively: (i) context coherence and retention, (ii) inter-component coordination and state management, (iii) tool use accuracy and selective execution, (iv) multi-step strategic planning, error detection, and recovery, and (v) real-time dynamic responsiveness. Our results show that while some architectures natively exhibit subsets of these properties, targeted augmentations substantially improve modular agent performance, especially in complex, multi-step, and real-time penetration testing tasks. |
| title | From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing |
| topic | Artificial Intelligence Computation and Language Machine Learning 68T50 I.2.7; I.5.4 |
| url | https://arxiv.org/abs/2509.14289 |