Optimized Deep Learning Models for Malware Detection under Concept Drift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Maillet, William, Marais, Benjamin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913454365868032
author Maillet, William
Marais, Benjamin
author_facet Maillet, William
Marais, Benjamin
contents Despite the promising results of machine learning models in malicious files detection, they face the problem of concept drift due to their constant evolution. This leads to declining performance over time, as the data distribution of the new files differs from the training one, requiring frequent model update. In this work, we propose a model-agnostic protocol to improve a baseline neural network against drift. We show the importance of feature reduction and training with the most recent validation set possible, and propose a loss function named Drift-Resilient Binary Cross-Entropy, an improvement to the classical Binary Cross-Entropy more effective against drift. We train our model on the EMBER dataset, published in2018, and evaluate it on a dataset of recent malicious files, collected between 2020 and 2023. Our improved model shows promising results, detecting 15.2% more malware than a baseline model.
format Preprint
id arxiv_https___arxiv_org_abs_2308_10821
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Optimized Deep Learning Models for Malware Detection under Concept Drift
Maillet, William
Marais, Benjamin
Cryptography and Security
Artificial Intelligence
Machine Learning
Despite the promising results of machine learning models in malicious files detection, they face the problem of concept drift due to their constant evolution. This leads to declining performance over time, as the data distribution of the new files differs from the training one, requiring frequent model update. In this work, we propose a model-agnostic protocol to improve a baseline neural network against drift. We show the importance of feature reduction and training with the most recent validation set possible, and propose a loss function named Drift-Resilient Binary Cross-Entropy, an improvement to the classical Binary Cross-Entropy more effective against drift. We train our model on the EMBER dataset, published in2018, and evaluate it on a dataset of recent malicious files, collected between 2020 and 2023. Our improved model shows promising results, detecting 15.2% more malware than a baseline model.
title Optimized Deep Learning Models for Malware Detection under Concept Drift
topic Cryptography and Security
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2308.10821