From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
Fuente:
arXiv
Saved in:
| Main Authors: | Conde, Pedro, Branquinho, Henrique, Mazzone, Valerio, Mendes, Bruno, Baptista, André, Moniz, Nuno |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
by: Ginige, Yasod, et al.
Published: (2025)
by: Ginige, Yasod, et al.
Published: (2025)
PentestJudge: Judging Agent Behavior Against Operational Requirements
by: Caldwell, Shane, et al.
Published: (2025)
by: Caldwell, Shane, et al.
Published: (2025)
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
by: Henke, Julius
Published: (2025)
by: Henke, Julius
Published: (2025)
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
by: Nieponice, Tomas, et al.
Published: (2025)
by: Nieponice, Tomas, et al.
Published: (2025)
PentestMCP: A Toolkit for Agentic Penetration Testing
by: Ezetta, Zachary, et al.
Published: (2025)
by: Ezetta, Zachary, et al.
Published: (2025)
A Preliminary Study on Using Large Language Models in Software Pentesting
by: Shashwat, Kumar, et al.
Published: (2024)
by: Shashwat, Kumar, et al.
Published: (2024)
Evaluating Privilege Usage of Agents with Real-World Tools
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
Generative Artificial Intelligence-Supported Pentesting: A Comparison between Claude Opus, GPT-4, and Copilot
by: Martínez, Antonio López, et al.
Published: (2025)
by: Martínez, Antonio López, et al.
Published: (2025)
Hacking, The Lazy Way: LLM Augmented Pentesting
by: Goyal, Dhruva, et al.
Published: (2024)
by: Goyal, Dhruva, et al.
Published: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
by: Reworr, et al.
Published: (2024)
by: Reworr, et al.
Published: (2024)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories
by: Shen, Chihao, et al.
Published: (2025)
by: Shen, Chihao, et al.
Published: (2025)
AdInject: Real-World Black-Box Attacks on Web Agents via Advertising Delivery
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
Hybrid Machine Learning Models for Intrusion Detection in IoT: Leveraging a Real-World IoT Dataset
by: Akif, Md Ahnaf, et al.
Published: (2025)
by: Akif, Md Ahnaf, et al.
Published: (2025)
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
by: Yao, Hongwei, et al.
Published: (2026)
by: Yao, Hongwei, et al.
Published: (2026)
ClawTrap: A MITM-Based Red-Teaming Framework for Real-World OpenClaw Security Evaluation
by: Zhao, Haochen, et al.
Published: (2026)
by: Zhao, Haochen, et al.
Published: (2026)
LogicEval: A Systematic Framework for Evaluating Automated Repair Techniques for Logical Vulnerabilities in Real-World Software
by: Rashid, Syed Md Mukit, et al.
Published: (2026)
by: Rashid, Syed Md Mukit, et al.
Published: (2026)
Agent Control Protocol: Admission Control for Agent Actions
by: Fernandez, Marcelo
Published: (2026)
by: Fernandez, Marcelo
Published: (2026)
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
by: Lin, Justin W., et al.
Published: (2025)
by: Lin, Justin W., et al.
Published: (2025)
PentestAgent: Incorporating LLM Agents to Automated Penetration Testing
by: Shen, Xiangmin, et al.
Published: (2024)
by: Shen, Xiangmin, et al.
Published: (2024)
Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
Enhancing Privacy in Federated Learning: Secure Aggregation for Real-World Healthcare Applications
by: Taiello, Riccardo, et al.
Published: (2024)
by: Taiello, Riccardo, et al.
Published: (2024)
Mobile GUI Agents under Real-world Threats: Are We There Yet?
by: Liu, Guohong, et al.
Published: (2025)
by: Liu, Guohong, et al.
Published: (2025)
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
by: Yang, Chenglin
Published: (2026)
by: Yang, Chenglin
Published: (2026)
How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World
by: Panebianco, Francesco, et al.
Published: (2025)
by: Panebianco, Francesco, et al.
Published: (2025)
Securing AI Agents with Information-Flow Control
by: Costa, Manuel, et al.
Published: (2025)
by: Costa, Manuel, et al.
Published: (2025)
Progent: Securing AI Agents with Privilege Control
by: Shi, Tianneng, et al.
Published: (2025)
by: Shi, Tianneng, et al.
Published: (2025)
MALCDF: A Distributed Multi-Agent LLM Framework for Real-Time Cyber
by: Bhardwaj, Arth, et al.
Published: (2025)
by: Bhardwaj, Arth, et al.
Published: (2025)
The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents
by: Wu, Baoyuan, et al.
Published: (2026)
by: Wu, Baoyuan, et al.
Published: (2026)
A Comparative Evaluation of AI Agent Security Guardrails
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
AIRGuard: Guarding Agent Actions with Runtime Authority Control
by: Qin, Suliu, et al.
Published: (2026)
by: Qin, Suliu, et al.
Published: (2026)
Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems
by: Chang, Hongyan, et al.
Published: (2026)
by: Chang, Hongyan, et al.
Published: (2026)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
by: Yang, Xianglin, et al.
Published: (2026)
by: Yang, Xianglin, et al.
Published: (2026)
OSS-CRS: Liberating AIxCC Cyber Reasoning Systems for Real-World Open-Source Security
by: Chin, Andrew, et al.
Published: (2026)
by: Chin, Andrew, et al.
Published: (2026)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
Federated Learning under Attack: Improving Gradient Inversion for Batch of Images
by: Leite, Luiz, et al.
Published: (2024)
by: Leite, Luiz, et al.
Published: (2024)
From Admission to Invariants: Measuring Deviation in Delegated Agent Systems
by: Fernandez, Marcelo
Published: (2026)
by: Fernandez, Marcelo
Published: (2026)
AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
by: Ngong, Ivoline C., et al.
Published: (2026)
by: Ngong, Ivoline C., et al.
Published: (2026)
Similar Items
-
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
by: Ginige, Yasod, et al.
Published: (2025) -
PentestJudge: Judging Agent Behavior Against Operational Requirements
by: Caldwell, Shane, et al.
Published: (2025) -
AutoPentest: Enhancing Vulnerability Management With Autonomous LLM Agents
by: Henke, Julius
Published: (2025) -
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
by: Nieponice, Tomas, et al.
Published: (2025) -
PentestMCP: A Toolkit for Agentic Penetration Testing
by: Ezetta, Zachary, et al.
Published: (2025)