Gray-Box Poisoning of Continuous Malware Ingestion Pipelines

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dolejš, Jan, Jureček, Martin, Lórencz, Róbert
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914534337282048
author Dolejš, Jan
Jureček, Martin
Lórencz, Róbert
author_facet Dolejš, Jan
Jureček, Martin
Lórencz, Róbert
contents Modern malware detection pipelines rely on continuous data ingestion and machine learning to counter the high volume of novel threats. This work investigates a realistic gray-box poisoning threat model targeting these pipelines. Using the secml_malware framework, we generate problem-space adversarial binaries through functionality-preserving manipulations, specifically Import Address Table (IAT) and section injections. We evaluate the impact of these poisoned samples when ingested into a defender's training set for a LightGBM malware detection model. Our empirical results demonstrate that subtle IAT-based perturbations enable compact poisoning samples that significantly degrade detection recall. These findings illustrate the inherent challenge of developing low-visibility adversarial perturbations that maintain high poisoning efficacy within continuous learning systems. We further evaluate a defense mechanism based on a homogeneous ensemble, which successfully identifies and filters up to 95.6% of poisoning attempts while maintaining a high retention rate for legitimate data. These findings emphasize the necessity of robust pre-ingestion validation in production pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04698
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Gray-Box Poisoning of Continuous Malware Ingestion Pipelines
Dolejš, Jan
Jureček, Martin
Lórencz, Róbert
Cryptography and Security
Machine Learning
Modern malware detection pipelines rely on continuous data ingestion and machine learning to counter the high volume of novel threats. This work investigates a realistic gray-box poisoning threat model targeting these pipelines. Using the secml_malware framework, we generate problem-space adversarial binaries through functionality-preserving manipulations, specifically Import Address Table (IAT) and section injections. We evaluate the impact of these poisoned samples when ingested into a defender's training set for a LightGBM malware detection model. Our empirical results demonstrate that subtle IAT-based perturbations enable compact poisoning samples that significantly degrade detection recall. These findings illustrate the inherent challenge of developing low-visibility adversarial perturbations that maintain high poisoning efficacy within continuous learning systems. We further evaluate a defense mechanism based on a homogeneous ensemble, which successfully identifies and filters up to 95.6% of poisoning attempts while maintaining a high retention rate for legitimate data. These findings emphasize the necessity of robust pre-ingestion validation in production pipelines.
title Gray-Box Poisoning of Continuous Malware Ingestion Pipelines
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.04698