Shell or Nothing: Real-World Benchmarks and Memory-Activated Agents for Automated Penetration Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mai, Wuyuao, Hong, Geng, Liu, Qi, Chen, Jinsong, Dai, Jiarun, Pan, Xudong, Zhang, Yuan, Yang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
von: Mai, Wuyuao, et al.
Veröffentlicht: (2025)
von: Mai, Wuyuao, et al.
Veröffentlicht: (2025)
MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
von: Wu, Mengying, et al.
Veröffentlicht: (2025)
von: Wu, Mengying, et al.
Veröffentlicht: (2025)
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
von: Luo, Jiaqi, et al.
Veröffentlicht: (2026)
von: Luo, Jiaqi, et al.
Veröffentlicht: (2026)
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
von: Fan, Yihe, et al.
Veröffentlicht: (2026)
von: Fan, Yihe, et al.
Veröffentlicht: (2026)
When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
von: Wu, Xinyi, et al.
Veröffentlicht: (2026)
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
von: Nakatani, Sho
Veröffentlicht: (2025)
von: Nakatani, Sho
Veröffentlicht: (2025)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
von: Shen, Xiangmin, et al.
Veröffentlicht: (2024)
von: Shen, Xiangmin, et al.
Veröffentlicht: (2024)
Automated Penetration Testing with LLM Agents and Classical Planning
von: Wang, Lingzhi, et al.
Veröffentlicht: (2025)
von: Wang, Lingzhi, et al.
Veröffentlicht: (2025)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
von: Lin, Justin W., et al.
Veröffentlicht: (2025)
von: Lin, Justin W., et al.
Veröffentlicht: (2025)
Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements
von: Isozaki, Isamu, et al.
Veröffentlicht: (2024)
von: Isozaki, Isamu, et al.
Veröffentlicht: (2024)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
von: Gioacchini, Luca, et al.
Veröffentlicht: (2024)
APT-Agent: Automated Penetration Testing using Large Language Models
von: Li, William Guanting, et al.
Veröffentlicht: (2026)
von: Li, William Guanting, et al.
Veröffentlicht: (2026)
Automated Penetration Testing: Formalization and Realization
von: Skandylas, Charilaos, et al.
Veröffentlicht: (2024)
von: Skandylas, Charilaos, et al.
Veröffentlicht: (2024)
Ethics Statements in Autonomous Penetration-Testing Agent Research
von: Happe, Andreas, et al.
Veröffentlicht: (2025)
von: Happe, Andreas, et al.
Veröffentlicht: (2025)
Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search
von: Shen, Yulin, et al.
Veröffentlicht: (2026)
von: Shen, Yulin, et al.
Veröffentlicht: (2026)
Reinforcement Learning for Automated Cybersecurity Penetration Testing
von: López-Montero, Daniel, et al.
Veröffentlicht: (2025)
von: López-Montero, Daniel, et al.
Veröffentlicht: (2025)
What Makes a Good LLM Agent for Real-world Penetration Testing?
von: Deng, Gelei, et al.
Veröffentlicht: (2026)
von: Deng, Gelei, et al.
Veröffentlicht: (2026)
Feedback-Guided Extraction of Knowledge Base from Retrieval-Augmented LLM Applications
von: Jiang, Changyue, et al.
Veröffentlicht: (2024)
von: Jiang, Changyue, et al.
Veröffentlicht: (2024)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
von: Lee, Hwiwon, et al.
Veröffentlicht: (2025)
Multi-Agent Penetration Testing AI for the Web
von: David, Isaac, et al.
Veröffentlicht: (2025)
von: David, Isaac, et al.
Veröffentlicht: (2025)
On the Surprising Efficacy of LLMs for Penetration-Testing
von: Happe, Andreas, et al.
Veröffentlicht: (2025)
von: Happe, Andreas, et al.
Veröffentlicht: (2025)
AWE: Adaptive Agents for Dynamic Web Penetration Testing
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
von: Jaswal, Akshat Singh, et al.
Veröffentlicht: (2026)
Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing
von: Abdulzada, Farzana
Veröffentlicht: (2025)
von: Abdulzada, Farzana
Veröffentlicht: (2025)
BreachSeek: A Multi-Agent Automated Penetration Tester
von: Alshehri, Ibrahim, et al.
Veröffentlicht: (2024)
von: Alshehri, Ibrahim, et al.
Veröffentlicht: (2024)
Insider Threats Mitigation: Role of Penetration Testing
von: Chauhan, Krutarth
Veröffentlicht: (2024)
von: Chauhan, Krutarth
Veröffentlicht: (2024)
Lessons from Penetration Tests on Large-Scale Agent Systems
von: Eykholt, Kevin, et al.
Veröffentlicht: (2026)
von: Eykholt, Kevin, et al.
Veröffentlicht: (2026)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
von: Fu, Yuchuan, et al.
Veröffentlicht: (2025)
Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration
von: Wu, Mengying, et al.
Veröffentlicht: (2024)
von: Wu, Mengying, et al.
Veröffentlicht: (2024)
Advanced Penetration Testing for Enhancing 5G Security
von: Smith-Haynes, Shari-Ann
Veröffentlicht: (2024)
von: Smith-Haynes, Shari-Ann
Veröffentlicht: (2024)
Penetration Testing for System Security: Methods and Practical Approaches
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
von: Zhang, Wei, et al.
Veröffentlicht: (2025)
A Comprehensive Evaluation and Practice of System Penetration Testing
von: Zhang, Chunyi, et al.
Veröffentlicht: (2025)
von: Zhang, Chunyi, et al.
Veröffentlicht: (2025)
Large language model-powered AI systems achieve self-replication with no human intervention
von: Pan, Xudong, et al.
Veröffentlicht: (2025)
von: Pan, Xudong, et al.
Veröffentlicht: (2025)
Penetration Testing of 5G Core Network Web Technologies
von: Giambartolomei, Filippo, et al.
Veröffentlicht: (2024)
von: Giambartolomei, Filippo, et al.
Veröffentlicht: (2024)
Critical Infrastructure Security: Penetration Testing and Exploit Development Perspectives
von: Orleans-Bosomtwe, Papa Kobina
Veröffentlicht: (2024)
von: Orleans-Bosomtwe, Papa Kobina
Veröffentlicht: (2024)
PTHelper: An open source tool to support the Penetration Testing process
von: de Gracia, Jacobo Casado, et al.
Veröffentlicht: (2024)
von: de Gracia, Jacobo Casado, et al.
Veröffentlicht: (2024)
SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection
von: Feng, Yang, et al.
Veröffentlicht: (2025)
von: Feng, Yang, et al.
Veröffentlicht: (2025)
StruPhantom: Evolutionary Injection Attacks on Black-Box Tabular Agents Powered by Large Language Models
von: Feng, Yang, et al.
Veröffentlicht: (2025)
von: Feng, Yang, et al.
Veröffentlicht: (2025)
AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation
von: Li, Changyi, et al.
Veröffentlicht: (2026)
von: Li, Changyi, et al.
Veröffentlicht: (2026)
AutoPT: How Far Are We from the End2End Automated Web Penetration Testing?
von: Wu, Benlong, et al.
Veröffentlicht: (2024)
von: Wu, Benlong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
von: Mai, Wuyuao, et al.
Veröffentlicht: (2025) -
MCPZoo: A Large-Scale Dataset of Runnable Model Context Protocol Servers for AI Agent
von: Wu, Mengying, et al.
Veröffentlicht: (2025) -
WebTrap Park: An Automated Platform for Systematic Security Evaluation of Web Agents
von: Wu, Xinyi, et al.
Veröffentlicht: (2026) -
AgentGuard: An Attribute-Based Access Control Framework for Tool-Use LLM-Based Agent
von: Luo, Jiaqi, et al.
Veröffentlicht: (2026) -
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
von: Fan, Yihe, et al.
Veröffentlicht: (2026)