Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
Fuente:
arXiv
Saved in:
| Main Authors: | Happe, Andreas, Cito, Jürgen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026)
by: Probst, Benjamin, et al.
Published: (2026)
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards
by: Normann, Philipp, et al.
Published: (2026)
by: Normann, Philipp, et al.
Published: (2026)
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
by: Happe, Andreas, et al.
Published: (2023)
by: Happe, Andreas, et al.
Published: (2023)
Got Root? A Linux Priv-Esc Benchmark
by: Happe, Andreas, et al.
Published: (2024)
by: Happe, Andreas, et al.
Published: (2024)
Ethics Statements in Autonomous Penetration-Testing Agent Research
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
On the Surprising Efficacy of LLMs for Penetration-Testing
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Can LLMs Hack Enterprise Networks? -- Replicated Computational Results (RCR) Report
by: Happe, Andreas, et al.
Published: (2026)
by: Happe, Andreas, et al.
Published: (2026)
Offensive Security for AI Systems: Concepts, Practices, and Applications
by: Harguess, Josh, et al.
Published: (2025)
by: Harguess, Josh, et al.
Published: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security
by: Valencia, Leroy Jacob
Published: (2024)
by: Valencia, Leroy Jacob
Published: (2024)
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
by: Rani, Nanda, et al.
Published: (2026)
by: Rani, Nanda, et al.
Published: (2026)
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
by: Udeshi, Meet, et al.
Published: (2025)
by: Udeshi, Meet, et al.
Published: (2025)
Dynamic Risk Assessments for Offensive Cybersecurity Agents
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security
by: Chua, Gabriel
Published: (2025)
by: Chua, Gabriel
Published: (2025)
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities
by: Kouremetis, Michael, et al.
Published: (2025)
by: Kouremetis, Michael, et al.
Published: (2025)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
SafeGenBench: A Benchmark Framework for Security Vulnerability Detection in LLM-Generated Code
by: Li, Xinghang, et al.
Published: (2025)
by: Li, Xinghang, et al.
Published: (2025)
Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards
by: Madhavan, Keerthana, et al.
Published: (2025)
by: Madhavan, Keerthana, et al.
Published: (2025)
From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists
by: Cecconello, Stefano, et al.
Published: (2026)
by: Cecconello, Stefano, et al.
Published: (2026)
On the (In)Security of LLM App Stores
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions
by: Susan, Stefan-Claudiu, et al.
Published: (2026)
by: Susan, Stefan-Claudiu, et al.
Published: (2026)
Weaponizing Language Models for Cybersecurity Offensive Operations: Automating Vulnerability Assessment Report Validation; A Review Paper
by: Almuhaidib, Abdulrahman S, et al.
Published: (2025)
by: Almuhaidib, Abdulrahman S, et al.
Published: (2025)
A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems
by: Wu, Fangzhou, et al.
Published: (2024)
by: Wu, Fangzhou, et al.
Published: (2024)
SkillTester: Benchmarking Utility and Security of Agent Skills
by: Wang, Leye, et al.
Published: (2026)
by: Wang, Leye, et al.
Published: (2026)
LLM Agents Should Employ Security Principles
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
aiXamine: Simplified LLM Safety and Security
by: Deniz, Fatih, et al.
Published: (2025)
by: Deniz, Fatih, et al.
Published: (2025)
A Framework for Formalizing LLM Agent Security
by: Siu, Vincent, et al.
Published: (2026)
by: Siu, Vincent, et al.
Published: (2026)
Towards Optimal Agentic Architectures for Offensive Security Tasks
by: David, Isaac, et al.
Published: (2026)
by: David, Isaac, et al.
Published: (2026)
LProtector: An LLM-driven Vulnerability Detection System
by: Sheng, Ze, et al.
Published: (2024)
by: Sheng, Ze, et al.
Published: (2024)
Towards more Practical Threat Models in Artificial Intelligence Security
by: Grosse, Kathrin, et al.
Published: (2023)
by: Grosse, Kathrin, et al.
Published: (2023)
Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling
by: Llewellyn, Mary, et al.
Published: (2025)
by: Llewellyn, Mary, et al.
Published: (2025)
Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems
by: Sharma, Gauri, et al.
Published: (2025)
by: Sharma, Gauri, et al.
Published: (2025)
Security-aware Semantic-driven ISAC via Paired Adversarial Residual Networks
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Information Security Based on LLM Approaches: A Review
by: Gong, Chang, et al.
Published: (2025)
by: Gong, Chang, et al.
Published: (2025)
Similar Items
-
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026) -
Cochise: A Reference Harness for Autonomous Penetration Testing
by: Happe, Andreas, et al.
Published: (2026) -
Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards
by: Normann, Philipp, et al.
Published: (2026) -
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
by: Happe, Andreas, et al.
Published: (2023) -
Got Root? A Linux Priv-Esc Benchmark
by: Happe, Andreas, et al.
Published: (2024)