Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gonçalves, José, Silva, Miguel, Maia, Eva, Praça, Isabel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908355631513600
author Gonçalves, José
Silva, Miguel
Maia, Eva
Praça, Isabel
author_facet Gonçalves, José
Silva, Miguel
Maia, Eva
Praça, Isabel
contents The application of Artificial Intelligence has become a powerful approach to detecting software vulnerabilities. However, effective vulnerability detection relies on accurately capturing the semantic structure of code and its contextual relationships. Given that the same functionality can be implemented in various forms, a preprocessing tool that standardizes code representation is important. This tool must be efficient, adaptable across programming languages, and capable of supporting new transformations. To address this challenge, we build on the existing SCoPE framework and introduce SCoPE2, an enhanced version with improved performance. We compare both versions in terms of processing time and memory usage and evaluate their impact on a Large Language Model (LLM) for vulnerability detection. Our results show a 97.3\% reduction in processing time with SCoPE2, along with an improved F1-score for the LLM, solely due to the refined preprocessing approach.
format Preprint
id arxiv_https___arxiv_org_abs_2505_05600
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
Gonçalves, José
Silva, Miguel
Maia, Eva
Praça, Isabel
Software Engineering
Machine Learning
The application of Artificial Intelligence has become a powerful approach to detecting software vulnerabilities. However, effective vulnerability detection relies on accurately capturing the semantic structure of code and its contextual relationships. Given that the same functionality can be implemented in various forms, a preprocessing tool that standardizes code representation is important. This tool must be efficient, adaptable across programming languages, and capable of supporting new transformations. To address this challenge, we build on the existing SCoPE framework and introduce SCoPE2, an enhanced version with improved performance. We compare both versions in terms of processing time and memory usage and evaluate their impact on a Large Language Model (LLM) for vulnerability detection. Our results show a 97.3\% reduction in processing time with SCoPE2, along with an improved F1-score for the LLM, solely due to the refined preprocessing approach.
title Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
topic Software Engineering
Machine Learning
url https://arxiv.org/abs/2505.05600