Cochise: A Reference Harness for Autonomous Penetration Testing
Fuente:
arXiv
Saved in:
| Main Authors: | Happe, Andreas, Cito, Jürgen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ethics Statements in Autonomous Penetration-Testing Agent Research
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
On the Surprising Efficacy of LLMs for Penetration-Testing
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
by: Happe, Andreas, et al.
Published: (2025)
by: Happe, Andreas, et al.
Published: (2025)
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)
by: Peng, Jiaren, et al.
Published: (2026)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks
by: Happe, Andreas, et al.
Published: (2023)
by: Happe, Andreas, et al.
Published: (2023)
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
by: Probst, Benjamin, et al.
Published: (2026)
by: Probst, Benjamin, et al.
Published: (2026)
Harnessing the Power of LLMs in Source Code Vulnerability Detection
by: Mahyari, Andrew A
Published: (2024)
by: Mahyari, Andrew A
Published: (2024)
Harnessing Large Language Models for Software Vulnerability Detection: A Comprehensive Benchmarking Study
by: Tamberg, Karl, et al.
Published: (2024)
by: Tamberg, Karl, et al.
Published: (2024)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards
by: Normann, Philipp, et al.
Published: (2026)
by: Normann, Philipp, et al.
Published: (2026)
Automated Penetration Testing: Formalization and Realization
by: Skandylas, Charilaos, et al.
Published: (2024)
by: Skandylas, Charilaos, et al.
Published: (2024)
Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode
by: Ji, Zimo, et al.
Published: (2026)
by: Ji, Zimo, et al.
Published: (2026)
MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
by: Lu, Xiaoyue, et al.
Published: (2026)
by: Lu, Xiaoyue, et al.
Published: (2026)
Testing Storage-System Correctness: Challenges, Fuzzing Limitations, and AI-Augmented Opportunities
by: Wang, Ying, et al.
Published: (2026)
by: Wang, Ying, et al.
Published: (2026)
CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming Solutions
by: Shi, Jingwei, et al.
Published: (2026)
by: Shi, Jingwei, et al.
Published: (2026)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
by: Deng, Gelei, et al.
Published: (2023)
by: Deng, Gelei, et al.
Published: (2023)
What Makes a Good LLM Agent for Real-world Penetration Testing?
by: Deng, Gelei, et al.
Published: (2026)
by: Deng, Gelei, et al.
Published: (2026)
A sketch of an AI control safety case
by: Korbak, Tomek, et al.
Published: (2025)
by: Korbak, Tomek, et al.
Published: (2025)
Security of LLM-generated Code: A Comparative Analysis
by: Morkonda, Srivathsan G, et al.
Published: (2026)
by: Morkonda, Srivathsan G, et al.
Published: (2026)
Innamark: A Whitespace Replacement Information-Hiding Method
by: Hellmeier, Malte, et al.
Published: (2025)
by: Hellmeier, Malte, et al.
Published: (2025)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
Graph Neural Networks for Vulnerability Detection: A Counterfactual Explanation
by: Chu, Zhaoyang, et al.
Published: (2024)
by: Chu, Zhaoyang, et al.
Published: (2024)
On the Security Risks of ML-based Malware Detection Systems: A Survey
by: He, Ping, et al.
Published: (2025)
by: He, Ping, et al.
Published: (2025)
Got Root? A Linux Priv-Esc Benchmark
by: Happe, Andreas, et al.
Published: (2024)
by: Happe, Andreas, et al.
Published: (2024)
A Comprehensive Study of Exploitable Patterns in Smart Contracts: From Vulnerability to Defense
by: Ding, Yuchen, et al.
Published: (2025)
by: Ding, Yuchen, et al.
Published: (2025)
Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code
by: Blain, Dominik, et al.
Published: (2026)
by: Blain, Dominik, et al.
Published: (2026)
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
by: Tracy, Tyler, et al.
Published: (2026)
by: Tracy, Tyler, et al.
Published: (2026)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
by: Yang, Hua, et al.
Published: (2025)
by: Yang, Hua, et al.
Published: (2025)
A Qualitative Study on Using ChatGPT for Software Security: Perception vs. Practicality
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
by: Kholoosi, M. Mehdi, et al.
Published: (2024)
Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6
by: Lorenzo, Luis Guzmán
Published: (2026)
by: Lorenzo, Luis Guzmán
Published: (2026)
SecureFixAgent: A Hybrid LLM Agent for Automated Python Static Vulnerability Repair
by: Gajjar, Jugal, et al.
Published: (2025)
by: Gajjar, Jugal, et al.
Published: (2025)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
by: Chen, Zhiyang, et al.
Published: (2025)
by: Chen, Zhiyang, et al.
Published: (2025)
CrossCert: A Cross-Checking Detection Approach to Patch Robustness Certification for Deep Learning Models
by: Zhou, Qilin, et al.
Published: (2024)
by: Zhou, Qilin, et al.
Published: (2024)
Leveraging Large Language Models for Cybersecurity Risk Assessment -- A Case from Forestry Cyber-Physical Systems
by: Gultekin, Fikret Mert, et al.
Published: (2025)
by: Gultekin, Fikret Mert, et al.
Published: (2025)
Deconstructing Obfuscation: A four-dimensional framework for evaluating Large Language Models assembly code deobfuscation capabilities
by: Tkachenko, Anton, et al.
Published: (2025)
by: Tkachenko, Anton, et al.
Published: (2025)
Similar Items
-
Ethics Statements in Autonomous Penetration-Testing Agent Research
by: Happe, Andreas, et al.
Published: (2025) -
Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks
by: Happe, Andreas, et al.
Published: (2025) -
On the Surprising Efficacy of LLMs for Penetration-Testing
by: Happe, Andreas, et al.
Published: (2025) -
Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design
by: Happe, Andreas, et al.
Published: (2025) -
Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing
by: Peng, Jiaren, et al.
Published: (2026)