Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ristea, Dan, McFadden, Shae, Shereen, Ezzeldin, Dwyer, Madeleine, Vyas, Sanyam, Hicks, Chris, Mavroudis, Vasilios
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909023873269760
author Ristea, Dan
McFadden, Shae
Shereen, Ezzeldin
Dwyer, Madeleine
Vyas, Sanyam
Hicks, Chris
Mavroudis, Vasilios
author_facet Ristea, Dan
McFadden, Shae
Shereen, Ezzeldin
Dwyer, Madeleine
Vyas, Sanyam
Hicks, Chris
Mavroudis, Vasilios
contents Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, especially as agentic coding frameworks increase the rate of code production. Over the last decade, a large body of research has applied machine learning machine learning to automate vulnerability detection (ML4AVD), yet self-reported performance on the most popular datasets shows no clear upward trend. The ML4AVD research community has identified several flaws in problem formulations, datasets, and metrics, but these are discussed in isolation, leaving the overarching problems that generate and reinforce these flaws unaddressed. We first systematize the field through a survey of 87 influential works based on their problem formulation, input and detection granularity, target programming languages, evaluation metrics, datasets, and detection approach. Drawing on this corpus and prior empirical work, we identify twelve pain points spanning the ML4AVD pipeline and show that they are self-reinforcing and causally inter-meshed: feedback loops between datasets, formulations, baselines, and metrics perpetuate each other and explain the field's persistent concentration on binary classification of C/C++ vulnerabilities at the function level. Thus, the field optimizes for a narrow and artificial problem that omits vulnerability type prediction, broader language support, and separation of input from detection granularity. We pair each pain point with concrete recommendations to break these loops. Finally, we use AIxCC as a case study to assess how well a recent high-profile effort aligns with these recommendations and reflect on the relevance of ML4AVD in the era of agentic AI.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11194
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
Ristea, Dan
McFadden, Shae
Shereen, Ezzeldin
Dwyer, Madeleine
Vyas, Sanyam
Hicks, Chris
Mavroudis, Vasilios
Software Engineering
Artificial Intelligence
Security vulnerabilities in software can have severe consequences; however, manual vulnerability detection is costly and does not scale, especially as agentic coding frameworks increase the rate of code production. Over the last decade, a large body of research has applied machine learning machine learning to automate vulnerability detection (ML4AVD), yet self-reported performance on the most popular datasets shows no clear upward trend. The ML4AVD research community has identified several flaws in problem formulations, datasets, and metrics, but these are discussed in isolation, leaving the overarching problems that generate and reinforce these flaws unaddressed. We first systematize the field through a survey of 87 influential works based on their problem formulation, input and detection granularity, target programming languages, evaluation metrics, datasets, and detection approach. Drawing on this corpus and prior empirical work, we identify twelve pain points spanning the ML4AVD pipeline and show that they are self-reinforcing and causally inter-meshed: feedback loops between datasets, formulations, baselines, and metrics perpetuate each other and explain the field's persistent concentration on binary classification of C/C++ vulnerabilities at the function level. Thus, the field optimizes for a narrow and artificial problem that omits vulnerability type prediction, broader language support, and separation of input from detection granularity. We pair each pain point with concrete recommendations to break these loops. Finally, we use AIxCC as a case study to assess how well a recent high-profile effort aligns with these recommendations and reflect on the relevance of ML4AVD in the era of agentic AI.
title Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points
topic Software Engineering
Artificial Intelligence
url https://arxiv.org/abs/2412.11194