How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency
Fuente:
arXiv
Guardado en:
| Autor principal: | Erdem, Galip Tolga |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
por: Afane, Khalifa, et al.
Publicado: (2024)
por: Afane, Khalifa, et al.
Publicado: (2024)
Multi-Agent Penetration Testing AI for the Web
por: David, Isaac, et al.
Publicado: (2025)
por: David, Isaac, et al.
Publicado: (2025)
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
por: Isozaki, Isamu, et al.
Publicado: (2024)
por: Isozaki, Isamu, et al.
Publicado: (2024)
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
por: Nakatani, Sho
Publicado: (2025)
por: Nakatani, Sho
Publicado: (2025)
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
por: Nguyen, Viet K., et al.
Publicado: (2025)
por: Nguyen, Viet K., et al.
Publicado: (2025)
How Effective Are Neural Networks for Fixing Security Vulnerabilities
por: Wu, Yi, et al.
Publicado: (2023)
por: Wu, Yi, et al.
Publicado: (2023)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
por: López-Montero, Daniel, et al.
Publicado: (2025)
por: López-Montero, Daniel, et al.
Publicado: (2025)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
por: Wu, Benlong, et al.
Publicado: (2024)
por: Wu, Benlong, et al.
Publicado: (2024)
Vulnerabilities in AI Code Generators: Exploring Targeted Data Poisoning Attacks
por: Cotroneo, Domenico, et al.
Publicado: (2023)
por: Cotroneo, Domenico, et al.
Publicado: (2023)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
por: Jaswal, Akshat Singh, et al.
Publicado: (2026)
por: Jaswal, Akshat Singh, et al.
Publicado: (2026)
PentestMCP: A Toolkit for Agentic Penetration Testing
por: Ezetta, Zachary, et al.
Publicado: (2025)
por: Ezetta, Zachary, et al.
Publicado: (2025)
Tailored Prompts, Targeted Protection: Vulnerability-Specific LLM Analysis for Smart Contracts
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning
por: Draganov, Andrew, et al.
Publicado: (2026)
por: Draganov, Andrew, et al.
Publicado: (2026)
Taught by the Flawed: How Dataset Insecurity Breeds Vulnerable AI Code
por: Xia, Catherine, et al.
Publicado: (2025)
por: Xia, Catherine, et al.
Publicado: (2025)
Lessons from Penetration Tests on Large-Scale Agent Systems
por: Eykholt, Kevin, et al.
Publicado: (2026)
por: Eykholt, Kevin, et al.
Publicado: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
por: Gioacchini, Luca, et al.
Publicado: (2024)
por: Gioacchini, Luca, et al.
Publicado: (2024)
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
por: Ferrag, Mohamed Amine, et al.
Publicado: (2024)
por: Ferrag, Mohamed Amine, et al.
Publicado: (2024)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
por: Quan, Haowei, et al.
Publicado: (2025)
por: Quan, Haowei, et al.
Publicado: (2025)
APT-Agent: Automated Penetration Testing using Large Language Models
por: Li, William Guanting, et al.
Publicado: (2026)
por: Li, William Guanting, et al.
Publicado: (2026)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
por: Lin, Justin W., et al.
Publicado: (2025)
por: Lin, Justin W., et al.
Publicado: (2025)
Leveraging AI to optimize website structure discovery during Penetration Testing
por: Antonelli, Diego, et al.
Publicado: (2021)
por: Antonelli, Diego, et al.
Publicado: (2021)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
por: Peng, Jiaren, et al.
Publicado: (2026)
por: Peng, Jiaren, et al.
Publicado: (2026)
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis
por: Ginige, Yasod, et al.
Publicado: (2026)
por: Ginige, Yasod, et al.
Publicado: (2026)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
por: Yang, Tiankai, et al.
Publicado: (2026)
por: Yang, Tiankai, et al.
Publicado: (2026)
What Breaks Embodied AI Security:LLM Vulnerabilities, CPS Flaws,or Something Else?
por: Ma, Boyang, et al.
Publicado: (2026)
por: Ma, Boyang, et al.
Publicado: (2026)
An Empirical Study of Vulnerability Detection using Federated Learning
por: Zhou, Peiheng, et al.
Publicado: (2024)
por: Zhou, Peiheng, et al.
Publicado: (2024)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
por: Yang, Ruozhao, et al.
Publicado: (2025)
por: Yang, Ruozhao, et al.
Publicado: (2025)
The System Prompt Is the Attack Surface: How LLM Agent Configuration Shapes Security and Creates Exploitable Vulnerabilities
por: Litvak, Ron
Publicado: (2026)
por: Litvak, Ron
Publicado: (2026)
How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition
por: Dziemian, Mateusz, et al.
Publicado: (2026)
por: Dziemian, Mateusz, et al.
Publicado: (2026)
Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks
por: Pasquini, Dario, et al.
Publicado: (2024)
por: Pasquini, Dario, et al.
Publicado: (2024)
Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms
por: Gasmi, Tarek, et al.
Publicado: (2025)
por: Gasmi, Tarek, et al.
Publicado: (2025)
Uplifted Attackers, Human Defenders: The Cyber Offense-Defense Balance for Trailing-Edge Organizations
por: Murphy, Benjamin, et al.
Publicado: (2025)
por: Murphy, Benjamin, et al.
Publicado: (2025)
DeSIA: Attribute Inference Attacks Against Limited Fixed Aggregate Statistics
por: Mao, Yifeng, et al.
Publicado: (2025)
por: Mao, Yifeng, et al.
Publicado: (2025)
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
por: Chua, Gabriel
Publicado: (2025)
por: Chua, Gabriel
Publicado: (2025)
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
por: Wang, Junran, et al.
Publicado: (2026)
por: Wang, Junran, et al.
Publicado: (2026)
AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks
por: Xu, Jiacen, et al.
Publicado: (2024)
por: Xu, Jiacen, et al.
Publicado: (2024)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
por: Gajjar, Jugal, et al.
Publicado: (2025)
por: Gajjar, Jugal, et al.
Publicado: (2025)
LProtector: An LLM-driven Vulnerability Detection System
por: Sheng, Ze, et al.
Publicado: (2024)
por: Sheng, Ze, et al.
Publicado: (2024)
Life-Cycle Routing Vulnerabilities of LLM Router
por: Lin, Qiqi, et al.
Publicado: (2025)
por: Lin, Qiqi, et al.
Publicado: (2025)
Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study
por: Chen, Zhihao, et al.
Publicado: (2026)
por: Chen, Zhihao, et al.
Publicado: (2026)
Ejemplares similares
-
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
por: Afane, Khalifa, et al.
Publicado: (2024) -
Multi-Agent Penetration Testing AI for the Web
por: David, Isaac, et al.
Publicado: (2025) -
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
por: Isozaki, Isamu, et al.
Publicado: (2024) -
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
por: Nakatani, Sho
Publicado: (2025) -
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
por: Nguyen, Viet K., et al.
Publicado: (2025)