Trace Gadgets: Minimizing Code Context for Machine Learning-Based Vulnerability Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mächtle, Felix, Loose, Nils, Schulz, Tim, Sieck, Florian, Serr, Jan-Niclas, Möller, Ralf, Eisenbarth, Thomas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915679355011072
author Mächtle, Felix
Loose, Nils
Schulz, Tim
Sieck, Florian
Serr, Jan-Niclas
Möller, Ralf
Eisenbarth, Thomas
author_facet Mächtle, Felix
Loose, Nils
Schulz, Tim
Sieck, Florian
Serr, Jan-Niclas
Möller, Ralf
Eisenbarth, Thomas
contents As the number of web applications and API endpoints exposed to the Internet continues to grow, so does the number of exploitable vulnerabilities. Manually identifying such vulnerabilities is tedious. Meanwhile, static security scanners tend to produce many false positives. While machine learning-based approaches are promising, they typically perform well only in scenarios where training and test data are closely related. A key challenge for ML-based vulnerability detection is providing suitable and concise code context, as excessively long contexts negatively affect the code comprehension capabilities of machine learning models, particularly smaller ones. This work introduces Trace Gadgets, a novel code representation that minimizes code context by removing non-related code. Trace Gadgets precisely capture the statements that cover the path to the vulnerability. As input for ML models, Trace Gadgets provide a minimal but complete context, thereby improving the detection performance. Moreover, we collect a large-scale dataset generated from real-world applications with manually curated labels to further improve the performance of ML-based vulnerability detectors. Our results show that state-of-the-art machine learning models perform best when using Trace Gadgets compared to previous code representations, surpassing the detection capabilities of industry-standard static scanners such as GitHub's CodeQL by at least 4% on a fully unseen dataset. By applying our framework to real-world applications, we identify and report previously unknown vulnerabilities in widely deployed software.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Trace Gadgets: Minimizing Code Context for Machine Learning-Based Vulnerability Prediction
Mächtle, Felix
Loose, Nils
Schulz, Tim
Sieck, Florian
Serr, Jan-Niclas
Möller, Ralf
Eisenbarth, Thomas
Cryptography and Security
Artificial Intelligence
As the number of web applications and API endpoints exposed to the Internet continues to grow, so does the number of exploitable vulnerabilities. Manually identifying such vulnerabilities is tedious. Meanwhile, static security scanners tend to produce many false positives. While machine learning-based approaches are promising, they typically perform well only in scenarios where training and test data are closely related. A key challenge for ML-based vulnerability detection is providing suitable and concise code context, as excessively long contexts negatively affect the code comprehension capabilities of machine learning models, particularly smaller ones. This work introduces Trace Gadgets, a novel code representation that minimizes code context by removing non-related code. Trace Gadgets precisely capture the statements that cover the path to the vulnerability. As input for ML models, Trace Gadgets provide a minimal but complete context, thereby improving the detection performance. Moreover, we collect a large-scale dataset generated from real-world applications with manually curated labels to further improve the performance of ML-based vulnerability detectors. Our results show that state-of-the-art machine learning models perform best when using Trace Gadgets compared to previous code representations, surpassing the detection capabilities of industry-standard static scanners such as GitHub's CodeQL by at least 4% on a fully unseen dataset. By applying our framework to real-world applications, we identify and report previously unknown vulnerabilities in widely deployed software.
title Trace Gadgets: Minimizing Code Context for Machine Learning-Based Vulnerability Prediction
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2504.13676