CTFusion: A CTF-based Benchmark for LLM Agent Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dongjun, Bae, Ga-eun, Yun, Insu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025)
by: Lee, Hwiwon, et al.
Published: (2025)
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)
by: Debenedetti, Edoardo, et al.
Published: (2024)
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
by: Rani, Nanda, et al.
Published: (2026)
by: Rani, Nanda, et al.
Published: (2026)
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
by: Lee, Hwiwon, et al.
Published: (2026)
by: Lee, Hwiwon, et al.
Published: (2026)
Takedown: How It's Done in Modern Coding Agent Exploits
by: Lee, Eunkyu, et al.
Published: (2025)
by: Lee, Eunkyu, et al.
Published: (2025)
Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
by: Kara, Su, et al.
Published: (2025)
by: Kara, Su, et al.
Published: (2025)
Memory-Induced Tool-Drift in LLM Agents
by: Dabas, Mahavir, et al.
Published: (2026)
by: Dabas, Mahavir, et al.
Published: (2026)
OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection
by: Flynt, Jeffrey
Published: (2026)
by: Flynt, Jeffrey
Published: (2026)
EMBER2024 -- A Benchmark Dataset for Holistic Evaluation of Malware Classifiers
by: Joyce, Robert J., et al.
Published: (2025)
by: Joyce, Robert J., et al.
Published: (2025)
LLM Benchmark Datasets Should Be Contamination-Resistant
by: Al-Lawati, Ali, et al.
Published: (2026)
by: Al-Lawati, Ali, et al.
Published: (2026)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
by: Hossain, S M Asif, et al.
Published: (2025)
by: Hossain, S M Asif, et al.
Published: (2025)
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
by: Shao, Minghao, et al.
Published: (2025)
by: Shao, Minghao, et al.
Published: (2025)
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
by: Anurin, Andrey, et al.
Published: (2024)
by: Anurin, Andrey, et al.
Published: (2024)
Design Patterns for Securing LLM Agents against Prompt Injections
by: Beurer-Kellner, Luca, et al.
Published: (2025)
by: Beurer-Kellner, Luca, et al.
Published: (2025)
Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications
by: Shahin, Nourin, et al.
Published: (2026)
by: Shahin, Nourin, et al.
Published: (2026)
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns
by: Shen, Xinyue, et al.
Published: (2025)
by: Shen, Xinyue, et al.
Published: (2025)
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
by: Nie, Yuzhou, et al.
Published: (2024)
by: Nie, Yuzhou, et al.
Published: (2024)
Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents
by: Lukáš, Ondřej, et al.
Published: (2026)
by: Lukáš, Ondřej, et al.
Published: (2026)
ControlNET: A Firewall for RAG-based LLM System
by: Yao, Hongwei, et al.
Published: (2025)
by: Yao, Hongwei, et al.
Published: (2025)
AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System
by: Ahmed, Chuadhry Mujeeb
Published: (2025)
by: Ahmed, Chuadhry Mujeeb
Published: (2025)
PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior
by: Flemings, James, et al.
Published: (2026)
by: Flemings, James, et al.
Published: (2026)
Fidelity, Diversity, and Privacy: A Multi-Dimensional LLM Evaluation for Clinical Data Augmentation
by: Iglesias, Guillermo, et al.
Published: (2026)
by: Iglesias, Guillermo, et al.
Published: (2026)
A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework
by: Chu, Kexin
Published: (2026)
by: Chu, Kexin
Published: (2026)
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
by: Jiralerspong, Thomas, et al.
Published: (2026)
by: Jiralerspong, Thomas, et al.
Published: (2026)
Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents
by: Tran, Toan, et al.
Published: (2026)
by: Tran, Toan, et al.
Published: (2026)
Adaptive Attacks Break Defenses Against Indirect Prompt Injection Attacks on LLM Agents
by: Zhan, Qiusi, et al.
Published: (2025)
by: Zhan, Qiusi, et al.
Published: (2025)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)
by: Eiras, Francisco, et al.
Published: (2025)
CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering
by: Gaddi, Matilda, et al.
Published: (2026)
by: Gaddi, Matilda, et al.
Published: (2026)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
by: Zizzo, Giulio, et al.
Published: (2025)
by: Zizzo, Giulio, et al.
Published: (2025)
A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions
by: Yang, Peiyu, et al.
Published: (2024)
by: Yang, Peiyu, et al.
Published: (2024)
A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
Adaptive Probe-based Steering for Robust LLM Jailbreaking
by: Chen, Junxi, et al.
Published: (2026)
by: Chen, Junxi, et al.
Published: (2026)
N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
by: Lin, Zheyu, et al.
Published: (2025)
by: Lin, Zheyu, et al.
Published: (2025)
TFHE-Coder: Evaluating LLM-agentic Fully Homomorphic Encryption Code Generation
by: Kumar, Mayank, et al.
Published: (2025)
by: Kumar, Mayank, et al.
Published: (2025)
StealthCup: Realistic, Multi-Stage, Evasion-Focused CTF for Benchmarking IDS
by: Kern, Manuel, et al.
Published: (2025)
by: Kern, Manuel, et al.
Published: (2025)
Similar Items
-
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024) -
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
by: Zhuo, Terry Yue, et al.
Published: (2025) -
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
by: Lee, Hwiwon, et al.
Published: (2025) -
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification
by: Zhang, Boyang, et al.
Published: (2024) -
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
by: Debenedetti, Edoardo, et al.
Published: (2024)