The Impact of Train-Test Leakage on Machine Learning-based Android Malware Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Guojun, Caragea, Doina, Ou, Xinming, Roy, Sankardas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913965404061696
author Liu, Guojun
Caragea, Doina
Ou, Xinming
Roy, Sankardas
author_facet Liu, Guojun
Caragea, Doina
Ou, Xinming
Roy, Sankardas
contents When machine learning is used for Android malware detection, an app needs to be represented in a numerical format for training and testing. We identify a widespread occurrence of distinct Android apps that have identical or nearly identical app representations. In particular, among app samples in the testing dataset, there can be a significant percentage of apps that have an identical or nearly identical representation to an app in the training dataset. This will lead to a data leakage problem that inflates a machine learning model's performance as measured on the testing dataset. The data leakage not only could lead to overly optimistic perceptions on the machine learning models' ability to generalize beyond the data on which they are trained, in some cases it could also lead to qualitatively different conclusions being drawn from the research. We present two case studies to illustrate this impact. In the first case study, the data leakage inflated the performance results but did not impact the overall conclusions made by the researchers in a qualitative way. In the second case study, the data leakage problem would have led to qualitatively different conclusions being drawn from the research. We further examine the real-world impact of the data leakage by dissecting the capability of memorization and the capability of generalization of a machine learning model, and show that by removing leakage from testing data, the evaluation results better reflect the machine learning model's utility in real-world Android malware detection scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19364
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Impact of Train-Test Leakage on Machine Learning-based Android Malware Detection
Liu, Guojun
Caragea, Doina
Ou, Xinming
Roy, Sankardas
Cryptography and Security
When machine learning is used for Android malware detection, an app needs to be represented in a numerical format for training and testing. We identify a widespread occurrence of distinct Android apps that have identical or nearly identical app representations. In particular, among app samples in the testing dataset, there can be a significant percentage of apps that have an identical or nearly identical representation to an app in the training dataset. This will lead to a data leakage problem that inflates a machine learning model's performance as measured on the testing dataset. The data leakage not only could lead to overly optimistic perceptions on the machine learning models' ability to generalize beyond the data on which they are trained, in some cases it could also lead to qualitatively different conclusions being drawn from the research. We present two case studies to illustrate this impact. In the first case study, the data leakage inflated the performance results but did not impact the overall conclusions made by the researchers in a qualitative way. In the second case study, the data leakage problem would have led to qualitatively different conclusions being drawn from the research. We further examine the real-world impact of the data leakage by dissecting the capability of memorization and the capability of generalization of a machine learning model, and show that by removing leakage from testing data, the evaluation results better reflect the machine learning model's utility in real-world Android malware detection scenarios.
title The Impact of Train-Test Leakage on Machine Learning-based Android Malware Detection
topic Cryptography and Security
url https://arxiv.org/abs/2410.19364