Towards Supporting Penetration Testing Education with Large Language Models: an Evaluation and Comparison
Fuente:
arXiv
Saved in:
| Main Authors: | Nizon-Deladoeuille, Martin, Stefánsson, Brynjólfur, Neukirchen, Helmut, Welsh, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Trust in Authentication Methods for Icelandic Digital Public Services
by: Stefánsson, Brynjólfur, et al.
Published: (2025)
by: Stefánsson, Brynjólfur, et al.
Published: (2025)
Towards Socio-Technical Topology-Aware Adaptive Threat Detection in Software Supply Chains
by: Welsh, Thomas, et al.
Published: (2025)
by: Welsh, Thomas, et al.
Published: (2025)
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
by: Adebimpe, Adetayo, et al.
Published: (2025)
by: Adebimpe, Adetayo, et al.
Published: (2025)
APT-Agent: Automated Penetration Testing using Large Language Models
by: Li, William Guanting, et al.
Published: (2026)
by: Li, William Guanting, et al.
Published: (2026)
A Comprehensive Evaluation and Practice of System Penetration Testing
by: Zhang, Chunyi, et al.
Published: (2025)
by: Zhang, Chunyi, et al.
Published: (2025)
On the Surprising Efficacy of LLMs for Penetration-Testing
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models
by: Luong, Phung Duc, et al.
Published: (2025)
by: Luong, Phung Duc, et al.
Published: (2025)
Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing
by: Abdulzada, Farzana
Published: (2025)
by: Abdulzada, Farzana
Published: (2025)
Insider Threats Mitigation: Role of Penetration Testing
by: Chauhan, Krutarth
Published: (2024)
by: Chauhan, Krutarth
Published: (2024)
Ethics Statements in Autonomous Penetration-Testing Agent Research
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Penetration Testing for System Security: Methods and Practical Approaches
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Automated Penetration Testing with LLM Agents and Classical Planning
by: Wang, Lingzhi, et al.
Published: (2025)
by: Wang, Lingzhi, et al.
Published: (2025)
Advanced Penetration Testing for Enhancing 5G Security
by: Smith-Haynes, Shari-Ann
Published: (2024)
by: Smith-Haynes, Shari-Ann
Published: (2024)
Lessons from Penetration Tests on Large-Scale Agent Systems
by: Eykholt, Kevin, et al.
Published: (2026)
by: Eykholt, Kevin, et al.
Published: (2026)
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
by: Isozaki, Isamu, et al.
Published: (2024)
by: Isozaki, Isamu, et al.
Published: (2024)
Penetration Testing of 5G Core Network Web Technologies
by: Giambartolomei, Filippo, et al.
Published: (2024)
by: Giambartolomei, Filippo, et al.
Published: (2024)
Critical Infrastructure Security: Penetration Testing and Exploit Development Perspectives
by: Orleans-Bosomtwe, Papa Kobina
Published: (2024)
by: Orleans-Bosomtwe, Papa Kobina
Published: (2024)
PTHelper: An open source tool to support the Penetration Testing process
by: de Gracia, Jacobo Casado, et al.
Published: (2024)
by: de Gracia, Jacobo Casado, et al.
Published: (2024)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
by: Shen, Xiangmin, et al.
Published: (2024)
by: Shen, Xiangmin, et al.
Published: (2024)
Automated Penetration Testing: Formalization and Realization
by: Skandylas, Charilaos, et al.
Published: (2024)
by: Skandylas, Charilaos, et al.
Published: (2024)
Supporting Human Raters with the Detection of Harmful Content using Large Language Models
by: Thomas, Kurt, et al.
Published: (2024)
by: Thomas, Kurt, et al.
Published: (2024)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Towards Automated Pentesting with Large Language Models
by: Bessa, Ricardo, et al.
Published: (2026)
by: Bessa, Ricardo, et al.
Published: (2026)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
by: López-Montero, Daniel, et al.
Published: (2025)
by: López-Montero, Daniel, et al.
Published: (2025)
Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
by: Mai, Wuyuao, et al.
Published: (2025)
by: Mai, Wuyuao, et al.
Published: (2025)
ICSSPulse: A Modular LLM-Assisted Platform for Industrial Control System Penetration Testing
by: Takaronis, Michail, et al.
Published: (2026)
by: Takaronis, Michail, et al.
Published: (2026)
Mind the Gap: Towards Generalizable Autonomous Penetration Testing via Domain Randomization and Meta-Reinforcement Learning
by: Zhou, Shicheng, et al.
Published: (2024)
by: Zhou, Shicheng, et al.
Published: (2024)
Enhancing Security Testing Software for Systems that Cannot be Subjected to the Risks of Penetration Testing Through the Incorporation of Multi-threading and and Other Capabilities
by: Tassava, Matthew, et al.
Published: (2024)
by: Tassava, Matthew, et al.
Published: (2024)
PentestMCP: A Toolkit for Agentic Penetration Testing
by: Ezetta, Zachary, et al.
Published: (2025)
by: Ezetta, Zachary, et al.
Published: (2025)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
by: Jaswal, Akshat Singh, et al.
Published: (2026)
by: Jaswal, Akshat Singh, et al.
Published: (2026)
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Red-MIRROR: Agentic LLM-based Autonomous Penetration Testing with Reflective Verification and Knowledge-augmented Interaction
by: Khang, Tran Vy, et al.
Published: (2026)
by: Khang, Tran Vy, et al.
Published: (2026)
Evaluation of Reinforcement Learning for Autonomous Penetration Testing using A3C, Q-learning and DQN
by: Becker, Norman, et al.
Published: (2024)
by: Becker, Norman, et al.
Published: (2024)
Towards Action Hijacking of Large Language Model-based Agent
by: Zhang, Yuyang, et al.
Published: (2024)
by: Zhang, Yuyang, et al.
Published: (2024)
AICrypto: Evaluating Cryptography Capabilities of Large Language Models
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
by: Yuan, Xiaohan, et al.
Published: (2024)
by: Yuan, Xiaohan, et al.
Published: (2024)
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
by: Nguyen, Viet K., et al.
Published: (2025)
by: Nguyen, Viet K., et al.
Published: (2025)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
PortGPT: Towards Automated Backporting Using Large Language Models
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Similar Items
-
Understanding Trust in Authentication Methods for Icelandic Digital Public Services
by: Stefánsson, Brynjólfur, et al.
Published: (2025) -
Towards Socio-Technical Topology-Aware Adaptive Threat Detection in Software Supply Chains
by: Welsh, Thomas, et al.
Published: (2025) -
SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
by: Adebimpe, Adetayo, et al.
Published: (2025) -
APT-Agent: Automated Penetration Testing using Large Language Models
by: Li, William Guanting, et al.
Published: (2026) -
A Comprehensive Evaluation and Practice of System Penetration Testing
by: Zhang, Chunyi, et al.
Published: (2025)