Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuo, Terry Yue, Wang, Dingmin, Ding, Hantian, Kumar, Varun, Wang, Zijian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cyber-Zero: Training Cybersecurity Agents without Runtime
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models
by: Berabi, Berkay, et al.
Published: (2024)
by: Berabi, Berkay, et al.
Published: (2024)
Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection
by: Wang, Zhaohui Geoffrey
Published: (2026)
by: Wang, Zhaohui Geoffrey
Published: (2026)
ANVIL: Anomaly-based Vulnerability Identification without Labelled Training Data
by: Wang, Weizhou, et al.
Published: (2024)
by: Wang, Weizhou, et al.
Published: (2024)
On the Security Vulnerabilities of Text-to-SQL Models
by: Peng, Xutan, et al.
Published: (2022)
by: Peng, Xutan, et al.
Published: (2022)
Learning-based Models for Vulnerability Detection: An Extensive Study
by: Ni, Chao, et al.
Published: (2024)
by: Ni, Chao, et al.
Published: (2024)
Security Vulnerability Detection with Multitask Self-Instructed Fine-Tuning of Large Language Models
by: Yang, Aidan Z. H., et al.
Published: (2024)
by: Yang, Aidan Z. H., et al.
Published: (2024)
Revisiting Pre-trained Language Models for Vulnerability Detection
by: Li, Youpeng, et al.
Published: (2025)
by: Li, Youpeng, et al.
Published: (2025)
From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization
by: Xi, Haoran, et al.
Published: (2025)
by: Xi, Haoran, et al.
Published: (2025)
Beyond Fidelity: Explaining Vulnerability Localization of Learning-based Detectors
by: Cheng, Baijun, et al.
Published: (2024)
by: Cheng, Baijun, et al.
Published: (2024)
Planning-Aware Code Infilling via Horizon-Length Prediction
by: Ding, Yifeng, et al.
Published: (2024)
by: Ding, Yifeng, et al.
Published: (2024)
Large Language Models for Code: Security Hardening and Adversarial Testing
by: He, Jingxuan, et al.
Published: (2023)
by: He, Jingxuan, et al.
Published: (2023)
To Err is Machine: Vulnerability Detection Challenges LLM Reasoning
by: Steenhoek, Benjamin, et al.
Published: (2024)
by: Steenhoek, Benjamin, et al.
Published: (2024)
A Study on Mixup-Inspired Augmentation Methods for Software Vulnerability Detection
by: Daneshvar, Seyed Shayan, et al.
Published: (2025)
by: Daneshvar, Seyed Shayan, et al.
Published: (2025)
Are Latent Vulnerabilities Hidden Gems for Software Vulnerability Prediction? An Empirical Study
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
Efficient Software Vulnerability Detection Using Transformer-based Models
by: Shaik, Sameer, et al.
Published: (2026)
by: Shaik, Sameer, et al.
Published: (2026)
TOSSS: a CVE-based Software Security Benchmark for Large Language Models
by: Damie, Marc, et al.
Published: (2026)
by: Damie, Marc, et al.
Published: (2026)
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
by: Hubinger, Evan, et al.
Published: (2024)
by: Hubinger, Evan, et al.
Published: (2024)
ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants
by: Xu, Xiangzhe, et al.
Published: (2025)
by: Xu, Xiangzhe, et al.
Published: (2025)
Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model
by: M, Keerthi Kumar., et al.
Published: (2026)
by: M, Keerthi Kumar., et al.
Published: (2026)
Software Vulnerability Prediction in Low-Resource Languages: An Empirical Study of CodeBERT and ChatGPT
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
Automatic Data Labeling for Software Vulnerability Prediction Models: How Far Are We?
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection
by: Lu, Chaomeng, et al.
Published: (2025)
by: Lu, Chaomeng, et al.
Published: (2025)
MVD: A Multi-Lingual Software Vulnerability Detection Framework
by: Zhang, Boyu, et al.
Published: (2024)
by: Zhang, Boyu, et al.
Published: (2024)
MARGIN: Margin-Aware Regularized Geometry for Imbalanced Vulnerability Detection
by: Zhang, Yuteng, et al.
Published: (2026)
by: Zhang, Yuteng, et al.
Published: (2026)
LLM-based Vulnerability Discovery through the Lens of Code Metrics
by: Weissberg, Felix, et al.
Published: (2025)
by: Weissberg, Felix, et al.
Published: (2025)
Understanding the Effectiveness of Large Language Models in Detecting Security Vulnerabilities
by: Khare, Avishree, et al.
Published: (2023)
by: Khare, Avishree, et al.
Published: (2023)
A Manually-Curated Dataset of Fixes to Vulnerabilities of Open-Source Software
by: Ponta, Serena E., et al.
Published: (2019)
by: Ponta, Serena E., et al.
Published: (2019)
An Empirical Study of Vulnerabilities in Python Packages and Their Detection
by: Quan, Haowei, et al.
Published: (2025)
by: Quan, Haowei, et al.
Published: (2025)
On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection
by: Hannan, Md Abdul, et al.
Published: (2025)
by: Hannan, Md Abdul, et al.
Published: (2025)
Automated Mapping of Vulnerability Advisories onto their Fix Commits in Open Source Repositories
by: Hommersom, Daan, et al.
Published: (2021)
by: Hommersom, Daan, et al.
Published: (2021)
Mitigating Data Imbalance for Software Vulnerability Assessment: Does Data Augmentation Help?
by: Le, Triet H. M., et al.
Published: (2024)
by: Le, Triet H. M., et al.
Published: (2024)
PurpCode: Reasoning for Safer Code Generation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Code-Centric Detection of Vulnerability-Fixing Commits: A Unified Benchmark and Empirical Study
by: Loose, Nils, et al.
Published: (2026)
by: Loose, Nils, et al.
Published: (2026)
RuleForge: Automated Generation and Validation for Web Vulnerability Detection at Scale
by: Garg, Ayush, et al.
Published: (2026)
by: Garg, Ayush, et al.
Published: (2026)
deepSURF: Detecting Memory Safety Vulnerabilities in Rust Through Fuzzing LLM-Augmented Harnesses
by: Androutsopoulos, Georgios, et al.
Published: (2025)
by: Androutsopoulos, Georgios, et al.
Published: (2025)
Security Is Relative: Training-Free Vulnerability Detection via Multi-Agent Behavioral Contract Synthesis
by: Wang, Yongchao, et al.
Published: (2026)
by: Wang, Yongchao, et al.
Published: (2026)
Automated Code-centric Software Vulnerability Assessment: How Far Are We? An Empirical Study in C/C++
by: Nguyen, Anh The, et al.
Published: (2024)
by: Nguyen, Anh The, et al.
Published: (2024)
Enhancing Pre-Trained Language Models for Vulnerability Detection via Semantic-Preserving Data Augmentation
by: Qi, Weiliang, et al.
Published: (2024)
by: Qi, Weiliang, et al.
Published: (2024)
AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection
by: Wang, Peiran, et al.
Published: (2025)
by: Wang, Peiran, et al.
Published: (2025)
Similar Items
-
Cyber-Zero: Training Cybersecurity Agents without Runtime
by: Zhuo, Terry Yue, et al.
Published: (2025) -
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models
by: Berabi, Berkay, et al.
Published: (2024) -
Strategic Heterogeneous Multi-Agent Architecture for Cost-Effective Code Vulnerability Detection
by: Wang, Zhaohui Geoffrey
Published: (2026) -
ANVIL: Anomaly-based Vulnerability Identification without Labelled Training Data
by: Wang, Weizhou, et al.
Published: (2024) -
On the Security Vulnerabilities of Text-to-SQL Models
by: Peng, Xutan, et al.
Published: (2022)