SecVulEval: Benchmarking LLMs for Real-World C/C++ Vulnerability Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Md Basim Uddin, Harzevili, Nima Shiri, Shin, Jiho, Pham, Hung Viet, Wang, Song |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StaAgent: An Agentic Framework for Testing Static Analyzers
by: Nnorom, Elijah, et al.
Published: (2025)
by: Nnorom, Elijah, et al.
Published: (2025)
Retrieval-Augmented Test Generation: How Far Are We?
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Checker Bug Detection and Repair in Deep Learning Libraries
by: Harzevili, Nima Shiri, et al.
Published: (2024)
by: Harzevili, Nima Shiri, et al.
Published: (2024)
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
by: Ni, Chao, et al.
Published: (2024)
by: Ni, Chao, et al.
Published: (2024)
VulWeaver: Weaving Broken Semantics for Grounded Vulnerability Detection
by: Cao, Yiheng, et al.
Published: (2026)
by: Cao, Yiheng, et al.
Published: (2026)
VulAgent: Hypothesis-Validation based Multi-Agent Vulnerability Detection
by: Wang, Ziliang, et al.
Published: (2025)
by: Wang, Ziliang, et al.
Published: (2025)
VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications
by: Zhu, Hao, et al.
Published: (2025)
by: Zhu, Hao, et al.
Published: (2025)
UntrustVul: An Automated Approach for Identifying Untrustworthy Alerts in Vulnerability Detection Models
by: Tung, Lam Nguyen, et al.
Published: (2025)
by: Tung, Lam Nguyen, et al.
Published: (2025)
A Survey on Query-based API Recommendation
by: Wei, Moshi, et al.
Published: (2023)
by: Wei, Moshi, et al.
Published: (2023)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
CleanVul: Automatic Function-Level Vulnerability Detection in Code Commits Using LLM Heuristics
by: Li, Yikun, et al.
Published: (2024)
by: Li, Yikun, et al.
Published: (2024)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
ABTest: Behavior-Driven Testing for AI Coding Agents
by: Dai, Wuyang, et al.
Published: (2026)
by: Dai, Wuyang, et al.
Published: (2026)
VulStamp: Vulnerability Assessment using Large Language Model
by: Shen, Hao, et al.
Published: (2025)
by: Shen, Hao, et al.
Published: (2025)
VulCoCo: A Simple Yet Effective Method for Detecting Vulnerable Code Clones
by: Bui, Tan, et al.
Published: (2025)
by: Bui, Tan, et al.
Published: (2025)
VulZoo: A Comprehensive Vulnerability Intelligence Dataset
by: Ruan, Bonan, et al.
Published: (2024)
by: Ruan, Bonan, et al.
Published: (2024)
PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities
by: Wei, Zichao, et al.
Published: (2025)
by: Wei, Zichao, et al.
Published: (2025)
Deep Learning Aided Software Vulnerability Detection: A Survey
by: Uddin, Md Nizam, et al.
Published: (2025)
by: Uddin, Md Nizam, et al.
Published: (2025)
Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
by: Du, Xueying, et al.
Published: (2024)
by: Du, Xueying, et al.
Published: (2024)
VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
VulGuard: An Unified Tool for Evaluating Just-In-Time Vulnerability Prediction Models
by: Nguyen, Duong, et al.
Published: (2025)
by: Nguyen, Duong, et al.
Published: (2025)
Secret Leak Detection in Software Issue Reports using LLMs: A Comprehensive Evaluation
by: Ahmed, Sadif, et al.
Published: (2024)
by: Ahmed, Sadif, et al.
Published: (2024)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
ReposVul: A Repository-Level High-Quality Vulnerability Dataset
by: Wang, Xinchen, et al.
Published: (2024)
by: Wang, Xinchen, et al.
Published: (2024)
StagedVulBERT: Multi-Granular Vulnerability Detection with a Novel Pre-trained Code Model
by: Jiang, Yuan, et al.
Published: (2024)
by: Jiang, Yuan, et al.
Published: (2024)
VulReaD: Knowledge-Graph-guided Software Vulnerability Reasoning and Detection
by: Mukhtar, Samal, et al.
Published: (2026)
by: Mukhtar, Samal, et al.
Published: (2026)
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
by: Jain, Kush, et al.
Published: (2024)
by: Jain, Kush, et al.
Published: (2024)
VulKey: Automated Vulnerability Repair Guided by Domain-Specific Repair Patterns
by: Li, Jia, et al.
Published: (2026)
by: Li, Jia, et al.
Published: (2026)
Prompt Engineering or Fine-Tuning: An Empirical Assessment of LLMs for Code
by: Shin, Jiho, et al.
Published: (2023)
by: Shin, Jiho, et al.
Published: (2023)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
by: Tian, Bowen, et al.
Published: (2025)
by: Tian, Bowen, et al.
Published: (2025)
MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution
by: Wu, Zihan, et al.
Published: (2026)
by: Wu, Zihan, et al.
Published: (2026)
VulRTex: A Reasoning-Guided Approach to Identify Vulnerabilities from Rich-Text Issue Report
by: Jiang, Ziyou, et al.
Published: (2025)
by: Jiang, Ziyou, et al.
Published: (2025)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
by: Li, Chenxin, et al.
Published: (2026)
by: Li, Chenxin, et al.
Published: (2026)
SWE-Bench+: Enhanced Coding Benchmark for LLMs
by: Aleithan, Reem, et al.
Published: (2024)
by: Aleithan, Reem, et al.
Published: (2024)
ReVul-CoT: Towards Effective Software Vulnerability Assessment with Retrieval-Augmented Generation and Chain-of-Thought Prompting
by: Chen, Zhijie, et al.
Published: (2025)
by: Chen, Zhijie, et al.
Published: (2025)
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
by: Weyssow, Martin, et al.
Published: (2025)
by: Weyssow, Martin, et al.
Published: (2025)
On the Role of Fault Localization Context for LLM-Based Program Repair
by: Sepidband, Melika, et al.
Published: (2026)
by: Sepidband, Melika, et al.
Published: (2026)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
by: Du, Junjia, et al.
Published: (2025)
by: Du, Junjia, et al.
Published: (2025)
Defect Prediction with Content-based Features
by: Pham, Hung Viet, et al.
Published: (2024)
by: Pham, Hung Viet, et al.
Published: (2024)
Similar Items
-
StaAgent: An Agentic Framework for Testing Static Analyzers
by: Nnorom, Elijah, et al.
Published: (2025) -
Retrieval-Augmented Test Generation: How Far Are We?
by: Shin, Jiho, et al.
Published: (2024) -
Checker Bug Detection and Repair in Deep Learning Libraries
by: Harzevili, Nima Shiri, et al.
Published: (2024) -
VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024) -
MegaVul: A C/C++ Vulnerability Dataset with Comprehensive Code Representation
by: Ni, Chao, et al.
Published: (2024)