Impact of Data Snooping on Deep Learning Models for Locating Vulnerabilities in Lifted Code

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: McCully, Gary A., Hastings, John D., Xu, Shengjie
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908732101754880
author McCully, Gary A.
Hastings, John D.
Xu, Shengjie
author_facet McCully, Gary A.
Hastings, John D.
Xu, Shengjie
contents This study examines the impact of data snooping on neural networks used to detect vulnerabilities in lifted code, and builds on previous research that used word2vec and unidirectional and bidirectional transformer-based embeddings. The research specifically focuses on how model performance is affected when embedding models are trained with datasets, which include samples used for neural network training and validation. The results show that introducing data snooping did not significantly alter model performance, suggesting that data snooping had a minimal impact or that samples randomly dropped as part of the methodology contained hidden features critical to achieving optimal performance. In addition, the findings reinforce the conclusions of previous research, which found that models trained with GPT-2 embeddings consistently outperformed neural networks trained with other embeddings. The fact that this holds even when data snooping is introduced into the embedding model indicates GPT-2's robustness in representing complex code features, even under less-than-ideal conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2412_02048
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Impact of Data Snooping on Deep Learning Models for Locating Vulnerabilities in Lifted Code
McCully, Gary A.
Hastings, John D.
Xu, Shengjie
Cryptography and Security
Computation and Language
Machine Learning
Software Engineering
D.4.6; I.2.6; I.5.1
This study examines the impact of data snooping on neural networks used to detect vulnerabilities in lifted code, and builds on previous research that used word2vec and unidirectional and bidirectional transformer-based embeddings. The research specifically focuses on how model performance is affected when embedding models are trained with datasets, which include samples used for neural network training and validation. The results show that introducing data snooping did not significantly alter model performance, suggesting that data snooping had a minimal impact or that samples randomly dropped as part of the methodology contained hidden features critical to achieving optimal performance. In addition, the findings reinforce the conclusions of previous research, which found that models trained with GPT-2 embeddings consistently outperformed neural networks trained with other embeddings. The fact that this holds even when data snooping is introduced into the embedding model indicates GPT-2's robustness in representing complex code features, even under less-than-ideal conditions.
title Impact of Data Snooping on Deep Learning Models for Locating Vulnerabilities in Lifted Code
topic Cryptography and Security
Computation and Language
Machine Learning
Software Engineering
D.4.6; I.2.6; I.5.1
url https://arxiv.org/abs/2412.02048