AutoPenBench: Benchmarking Generative Agents for Penetration Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gioacchini, Luca, Mellia, Marco, Drago, Idilio, Delsanto, Alexander, Siracusano, Giuseppe, Bifulco, Roberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Agentic Honeynet Configuration
von: Mirra, Federico, et al.
Veröffentlicht: (2026)
von: Mirra, Federico, et al.
Veröffentlicht: (2026)
Autonomous LLM Agents & CTFs: A Second Look
von: Bouchari, Youness, et al.
Veröffentlicht: (2026)
von: Bouchari, Youness, et al.
Veröffentlicht: (2026)
Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning
von: Huang, Jianan, et al.
Veröffentlicht: (2026)
von: Huang, Jianan, et al.
Veröffentlicht: (2026)
LogPrécis: Unleashing Language Models for Automated Malicious Log Analysis
von: Boffa, Matteo, et al.
Veröffentlicht: (2023)
von: Boffa, Matteo, et al.
Veröffentlicht: (2023)
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
von: Nakatani, Sho
Veröffentlicht: (2025)
von: Nakatani, Sho
Veröffentlicht: (2025)
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis
von: Ginige, Yasod, et al.
Veröffentlicht: (2026)
von: Ginige, Yasod, et al.
Veröffentlicht: (2026)
Generic Multi-modal Representation Learning for Network Traffic Analysis
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
DeePen: Penetration Testing for Audio Deepfake Detection
von: Müller, Nicolas, et al.
Veröffentlicht: (2025)
von: Müller, Nicolas, et al.
Veröffentlicht: (2025)
Multi-Agent Penetration Testing AI for the Web
von: David, Isaac, et al.
Veröffentlicht: (2025)
von: David, Isaac, et al.
Veröffentlicht: (2025)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
Lessons from Penetration Tests on Large-Scale Agent Systems
von: Eykholt, Kevin, et al.
Veröffentlicht: (2026)
von: Eykholt, Kevin, et al.
Veröffentlicht: (2026)
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
von: Isozaki, Isamu, et al.
Veröffentlicht: (2024)
von: Isozaki, Isamu, et al.
Veröffentlicht: (2024)
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
von: Wu, Benlong, et al.
Veröffentlicht: (2024)
von: Wu, Benlong, et al.
Veröffentlicht: (2024)
PenTest++: Elevating Ethical Hacking with AI and Automation
von: Al-Sinani, Haitham S., et al.
Veröffentlicht: (2025)
von: Al-Sinani, Haitham S., et al.
Veröffentlicht: (2025)
APT-Agent: Automated Penetration Testing using Large Language Models
von: Li, William Guanting, et al.
Veröffentlicht: (2026)
von: Li, William Guanting, et al.
Veröffentlicht: (2026)
Knowledge-Informed Auto-Penetration Testing Based on Reinforcement Learning with Reward Machine
von: Li, Yuanliang, et al.
Veröffentlicht: (2024)
von: Li, Yuanliang, et al.
Veröffentlicht: (2024)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
von: López-Montero, Daniel, et al.
Veröffentlicht: (2025)
von: López-Montero, Daniel, et al.
Veröffentlicht: (2025)
xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models
von: Luong, Phung Duc, et al.
Veröffentlicht: (2025)
von: Luong, Phung Duc, et al.
Veröffentlicht: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
von: Zhang, Hanrong, et al.
Veröffentlicht: (2024)
PentestMCP: A Toolkit for Agentic Penetration Testing
von: Ezetta, Zachary, et al.
Veröffentlicht: (2025)
von: Ezetta, Zachary, et al.
Veröffentlicht: (2025)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
von: Lee, Seunghyun, et al.
Veröffentlicht: (2026)
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
von: Najt, Elle, et al.
Veröffentlicht: (2026)
von: Najt, Elle, et al.
Veröffentlicht: (2026)
ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense
von: Lau, Nancy, et al.
Veröffentlicht: (2026)
von: Lau, Nancy, et al.
Veröffentlicht: (2026)
CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios
von: Lim, Taein, et al.
Veröffentlicht: (2026)
von: Lim, Taein, et al.
Veröffentlicht: (2026)
MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
von: Yang, Yixuan, et al.
Veröffentlicht: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
von: Lin, Justin W., et al.
Veröffentlicht: (2025)
von: Lin, Justin W., et al.
Veröffentlicht: (2025)
GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization
von: Nimase, Ojas, et al.
Veröffentlicht: (2026)
von: Nimase, Ojas, et al.
Veröffentlicht: (2026)
BreachSeek: A Multi-Agent Automated Penetration Tester
von: Alshehri, Ibrahim, et al.
Veröffentlicht: (2024)
von: Alshehri, Ibrahim, et al.
Veröffentlicht: (2024)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
von: Carlini, Nicholas, et al.
Veröffentlicht: (2025)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
von: Yang, Ruozhao, et al.
Veröffentlicht: (2025)
von: Yang, Ruozhao, et al.
Veröffentlicht: (2025)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
von: Shen, Chihao, et al.
Veröffentlicht: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
von: Zhang, Dongsen, et al.
Veröffentlicht: (2025)
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
von: Chen, Qirui, et al.
Veröffentlicht: (2026)
Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks
von: Nguyen, Viet K., et al.
Veröffentlicht: (2025)
von: Nguyen, Viet K., et al.
Veröffentlicht: (2025)
SastBench: A Benchmark for Testing Agentic SAST Triage
von: Feiglin, Jake, et al.
Veröffentlicht: (2026)
von: Feiglin, Jake, et al.
Veröffentlicht: (2026)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
von: Li, Xinghang, et al.
Veröffentlicht: (2025)
Incorporation of Verifier Functionality in the Software for Operations and Network Attack Results Review and the Autonomous Penetration Testing System
von: Milbrath, Jordan, et al.
Veröffentlicht: (2024)
von: Milbrath, Jordan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Agentic Honeynet Configuration
von: Mirra, Federico, et al.
Veröffentlicht: (2026) -
Autonomous LLM Agents & CTFs: A Second Look
von: Bouchari, Youness, et al.
Veröffentlicht: (2026) -
Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning
von: Huang, Jianan, et al.
Veröffentlicht: (2026) -
LogPrécis: Unleashing Language Models for Automated Malicious Log Analysis
von: Boffa, Matteo, et al.
Veröffentlicht: (2023) -
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
von: Nakatani, Sho
Veröffentlicht: (2025)