Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
Fuente:
arXiv
Saved in:
| Main Authors: | Isozaki, Isamu, Shrestha, Manil, Console, Rick, Kim, Edward |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Secure Multiparty Generative AI
by: Shrestha, Manil, et al.
Published: (2024)
by: Shrestha, Manil, et al.
Published: (2024)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
by: López-Montero, Daniel, et al.
Published: (2025)
by: López-Montero, Daniel, et al.
Published: (2025)
APT-Agent: Automated Penetration Testing using Large Language Models
by: Li, William Guanting, et al.
Published: (2026)
by: Li, William Guanting, et al.
Published: (2026)
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
by: Nakatani, Sho
Published: (2025)
by: Nakatani, Sho
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
by: Gioacchini, Luca, et al.
Published: (2024)
by: Gioacchini, Luca, et al.
Published: (2024)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
by: Wu, Benlong, et al.
Published: (2024)
by: Wu, Benlong, et al.
Published: (2024)
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis
by: Ginige, Yasod, et al.
Published: (2026)
by: Ginige, Yasod, et al.
Published: (2026)
Multi-Agent Penetration Testing AI for the Web
by: David, Isaac, et al.
Published: (2025)
by: David, Isaac, et al.
Published: (2025)
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
by: Nguyen, Viet K., et al.
Published: (2025)
by: Nguyen, Viet K., et al.
Published: (2025)
BreachSeek: A Multi-Agent Automated Penetration Tester
by: Alshehri, Ibrahim, et al.
Published: (2024)
by: Alshehri, Ibrahim, et al.
Published: (2024)
PentestMCP: A Toolkit for Agentic Penetration Testing
by: Ezetta, Zachary, et al.
Published: (2025)
by: Ezetta, Zachary, et al.
Published: (2025)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
by: Jaswal, Akshat Singh, et al.
Published: (2026)
by: Jaswal, Akshat Singh, et al.
Published: (2026)
Lessons from Penetration Tests on Large-Scale Agent Systems
by: Eykholt, Kevin, et al.
Published: (2026)
by: Eykholt, Kevin, et al.
Published: (2026)
Towards Automating Blockchain Consensus Verification with IsabeLLM
by: Jones, Elliot, et al.
Published: (2026)
by: Jones, Elliot, et al.
Published: (2026)
How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency
by: Erdem, Galip Tolga
Published: (2026)
by: Erdem, Galip Tolga
Published: (2026)
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
by: Lim, Taein, et al.
Published: (2026)
by: Lim, Taein, et al.
Published: (2026)
Incorporation of Verifier Functionality in the Software for Operations and Network Attack Results Review and the Autonomous Penetration Testing System
by: Milbrath, Jordan, et al.
Published: (2024)
by: Milbrath, Jordan, et al.
Published: (2024)
Evaluation of Reinforcement Learning for Autonomous Penetration Testing using A3C, Q-learning and DQN
by: Becker, Norman, et al.
Published: (2024)
by: Becker, Norman, et al.
Published: (2024)
xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models
by: Luong, Phung Duc, et al.
Published: (2025)
by: Luong, Phung Duc, et al.
Published: (2025)
AiRacleX: Automated Detection of Price Oracle Manipulations via LLM-Driven Knowledge Mining and Prompt Generation
by: Gao, Bo, et al.
Published: (2025)
by: Gao, Bo, et al.
Published: (2025)
CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher
by: Pratama, Derry, et al.
Published: (2024)
by: Pratama, Derry, et al.
Published: (2024)
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
Towards the Development of an LLM-Based Methodology for Automated Security Profiling in Compliance with Ukrainian Cybersecurity Regulations
by: Shafranskyi, Daniil, et al.
Published: (2026)
by: Shafranskyi, Daniil, et al.
Published: (2026)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
by: Huang, Enhao, et al.
Published: (2025)
by: Huang, Enhao, et al.
Published: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
Leveraging AI to optimize website structure discovery during Penetration Testing
by: Antonelli, Diego, et al.
Published: (2021)
by: Antonelli, Diego, et al.
Published: (2021)
ADAPT: A Game-Theoretic and Neuro-Symbolic Framework for Automated Distributed Adaptive Penetration Testing
by: Lei, Haozhe, et al.
Published: (2024)
by: Lei, Haozhe, et al.
Published: (2024)
PenTest++: Elevating Ethical Hacking with AI and Automation
by: Al-Sinani, Haitham S., et al.
Published: (2025)
by: Al-Sinani, Haitham S., et al.
Published: (2025)
Automated Penetration Testing with LLM Agents and Classical Planning
by: Wang, Lingzhi, et al.
Published: (2025)
by: Wang, Lingzhi, et al.
Published: (2025)
Introducing the Generative Application Firewall (GAF)
by: Farreny, Joan Vendrell, et al.
Published: (2026)
by: Farreny, Joan Vendrell, et al.
Published: (2026)
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine
by: Li, Yuanliang, et al.
Published: (2024)
by: Li, Yuanliang, et al.
Published: (2024)
LLM-Assisted Proactive Threat Intelligence for Automated Reasoning
by: Paul, Shuva, et al.
Published: (2025)
by: Paul, Shuva, et al.
Published: (2025)
Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions
by: Susan, Stefan-Claudiu, et al.
Published: (2026)
by: Susan, Stefan-Claudiu, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
by: Ginige, Yasod, et al.
Published: (2025)
by: Ginige, Yasod, et al.
Published: (2025)
HonestCyberEval: An AI Cyber Risk Benchmark for Automated Software Exploitation
by: Ristea, Dan, et al.
Published: (2024)
by: Ristea, Dan, et al.
Published: (2024)
Similar Items
-
Secure Multiparty Generative AI
by: Shrestha, Manil, et al.
Published: (2024) -
Reinforcement Learning for Automated Cybersecurity Penetration Testing
by: López-Montero, Daniel, et al.
Published: (2025) -
APT-Agent: Automated Penetration Testing using Large Language Models
by: Li, William Guanting, et al.
Published: (2026) -
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
by: Nakatani, Sho
Published: (2025) -
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)