PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ruozhao, Cheng, Mingfei, Deng, Gelei, Zhang, Tianwei, Wang, Junjie, Xie, Xiaofei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026)
by: Yang, Ruozhao, et al.
Published: (2026)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
Automated Penetration Testing: Formalization and Realization
by: Skandylas, Charilaos, et al.
Published: (2024)
by: Skandylas, Charilaos, et al.
Published: (2024)
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
PatchFuzz: Patch Fuzzing for JavaScript Engines
by: Wang, Junjie, et al.
Published: (2025)
by: Wang, Junjie, et al.
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
Towards Demystifying and Repairing LLM-in-the-Loop Vulnerabilities
by: Ma, Yujie, et al.
Published: (2026)
by: Ma, Yujie, et al.
Published: (2026)
SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysis
by: Lingxiang, Wang, et al.
Published: (2025)
by: Lingxiang, Wang, et al.
Published: (2025)
OpDiffer: LLM-Assisted Opcode-Level Differential Testing of Ethereum Virtual Machine
by: Ma, Jie, et al.
Published: (2025)
by: Ma, Jie, et al.
Published: (2025)
CAShift: Benchmarking Log-Based Cloud Attack Detection under Normality Shift
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Towards Understanding and Characterizing Vulnerabilities in Intelligent Connected Vehicles through Real-World Exploits
by: Wang, Yuelin, et al.
Published: (2026)
by: Wang, Yuelin, et al.
Published: (2026)
Probing Privacy Leaks in LLM-based Code Generation via Test Generation
by: Ge, Yifei, et al.
Published: (2026)
by: Ge, Yifei, et al.
Published: (2026)
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
by: Chu, Junjie, et al.
Published: (2026)
by: Chu, Junjie, et al.
Published: (2026)
FCGHunter: Towards Evaluating Robustness of Graph-Based Android Malware Detection
by: Song, Shiwen, et al.
Published: (2025)
by: Song, Shiwen, et al.
Published: (2025)
AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Prompt Injection attack against LLM-integrated Applications
by: Liu, Yi, et al.
Published: (2023)
by: Liu, Yi, et al.
Published: (2023)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Scalable Test Generation to Trigger Rare Targets in High-Level Synthesizable IPs for Cloud FPGAs
by: Debnath, Mukta, et al.
Published: (2024)
by: Debnath, Mukta, et al.
Published: (2024)
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
by: Hu, Qi, et al.
Published: (2026)
by: Hu, Qi, et al.
Published: (2026)
ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents
by: Wu, Jiangrong, et al.
Published: (2026)
by: Wu, Jiangrong, et al.
Published: (2026)
Understanding the Supply Chain and Risks of Large Language Model Applications
by: Ma, Yujie, et al.
Published: (2025)
by: Ma, Yujie, et al.
Published: (2025)
How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test
by: Tessa, Melissa, et al.
Published: (2026)
by: Tessa, Melissa, et al.
Published: (2026)
VERCATION: Precise Vulnerable Open-source Software Version Identification based on Static Analysis and LLM
by: Cheng, Yiran, et al.
Published: (2024)
by: Cheng, Yiran, et al.
Published: (2024)
Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test
by: Haque, Mohd Ariful, et al.
Published: (2026)
by: Haque, Mohd Ariful, et al.
Published: (2026)
Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety
by: Chen, Xuan, et al.
Published: (2026)
by: Chen, Xuan, et al.
Published: (2026)
AutoDFBench 1.0: A Benchmarking Framework for Digital Forensic Tool Testing and Generated Code Evaluation
by: Wickramasekara, Akila, et al.
Published: (2025)
by: Wickramasekara, Akila, et al.
Published: (2025)
VulnRepairEval: An Exploit-Based Evaluation Framework for Assessing Large Language Model Vulnerability Repair Capabilities
by: Wang, Weizhe, et al.
Published: (2025)
by: Wang, Weizhe, et al.
Published: (2025)
PrediQL: Automated Testing of GraphQL APIs with LLMs
by: Liu, Shaolun, et al.
Published: (2025)
by: Liu, Shaolun, et al.
Published: (2025)
SCRUTINEER: Detecting Logic-Level Usage Violations of Reusable Components in Smart Contracts
by: Lin, Xingshuang, et al.
Published: (2025)
by: Lin, Xingshuang, et al.
Published: (2025)
Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval
by: Sabir, Bushra, et al.
Published: (2026)
by: Sabir, Bushra, et al.
Published: (2026)
QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
by: Ameen, Mohd Ruhul, et al.
Published: (2026)
ReposVul: A Repository-Level High-Quality Vulnerability Dataset
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
Virtualization-based Penetration Testing Study for Detecting Accessibility Abuse Vulnerabilities in Banking Apps in East and Southeast Asia
by: Minn, Wei, et al.
Published: (2026)
by: Minn, Wei, et al.
Published: (2026)
Security Testing of RESTful APIs With Test Case Mutation
by: Salva, Sebastien, et al.
Published: (2024)
by: Salva, Sebastien, et al.
Published: (2024)
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
by: Wei, Zichao, et al.
Published: (2025)
by: Wei, Zichao, et al.
Published: (2025)
Similar Items
-
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026) -
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
by: Yang, Ruozhao, et al.
Published: (2026) -
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023) -
Automated Penetration Testing: Formalization and Realization
by: Skandylas, Charilaos, et al.
Published: (2024) -
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)