Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yildiz, Alperen, Teo, Sin G., Lou, Yiling, Feng, Yebo, Wang, Chong, Divakaran, Dinil M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915202828599296
author Yildiz, Alperen
Teo, Sin G.
Lou, Yiling
Feng, Yebo
Wang, Chong
Divakaran, Dinil M.
author_facet Yildiz, Alperen
Teo, Sin G.
Lou, Yiling
Feng, Yebo
Wang, Chong
Divakaran, Dinil M.
contents Large Language Models (LLMs) have shown promise in software vulnerability detection, particularly on function-level benchmarks like Devign and BigVul. However, real-world detection requires interprocedural analysis, as vulnerabilities often emerge through multi-hop function calls rather than isolated functions. While repository-level benchmarks like ReposVul and VulEval introduce interprocedural context, they remain computationally expensive, lack pairwise evaluation of vulnerability fixes, and explore limited context retrieval, limiting their practicality. We introduce JitVul, a JIT vulnerability detection benchmark linking each function to its vulnerability-introducing and fixing commits. Built from 879 CVEs spanning 91 vulnerability types, JitVul enables comprehensive evaluation of detection capabilities. Our results show that ReAct Agents, leveraging thought-action-observation and interprocedural context, perform better than LLMs in distinguishing vulnerable from benign code. While prompting strategies like Chain-of-Thought help LLMs, ReAct Agents require further refinement. Both methods show inconsistencies, either misidentifying vulnerabilities or over-analyzing security guards, indicating significant room for improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03586
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
Yildiz, Alperen
Teo, Sin G.
Lou, Yiling
Feng, Yebo
Wang, Chong
Divakaran, Dinil M.
Cryptography and Security
Large Language Models (LLMs) have shown promise in software vulnerability detection, particularly on function-level benchmarks like Devign and BigVul. However, real-world detection requires interprocedural analysis, as vulnerabilities often emerge through multi-hop function calls rather than isolated functions. While repository-level benchmarks like ReposVul and VulEval introduce interprocedural context, they remain computationally expensive, lack pairwise evaluation of vulnerability fixes, and explore limited context retrieval, limiting their practicality. We introduce JitVul, a JIT vulnerability detection benchmark linking each function to its vulnerability-introducing and fixing commits. Built from 879 CVEs spanning 91 vulnerability types, JitVul enables comprehensive evaluation of detection capabilities. Our results show that ReAct Agents, leveraging thought-action-observation and interprocedural context, perform better than LLMs in distinguishing vulnerable from benign code. While prompting strategies like Chain-of-Thought help LLMs, ReAct Agents require further refinement. Both methods show inconsistencies, either misidentifying vulnerabilities or over-analyzing security guards, indicating significant room for improvement.
title Benchmarking LLMs and LLM-based Agents in Practical Vulnerability Detection for Code Repositories
topic Cryptography and Security
url https://arxiv.org/abs/2503.03586