Detecting Concept Drift in Evolving Malware Families Using Rule-Based Classifier Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kalný, Tomáš, Jureček, Martin, Stamp, Mark
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913059242508288
author Kalný, Tomáš
Jureček, Martin
Stamp, Mark
author_facet Kalný, Tomáš
Jureček, Martin
Stamp, Mark
contents This work proposes a structural approach to concept drift detection in malware classification using decision tree rulesets. Classifiers are trained across temporal windows on the EMBER2024 dataset, and drift is quantified by comparing extracted rule representations using feature importance, prediction agreement, activation stability, and coverage metrics. These metrics are correlated with both accuracy degradation and data distribution shift as complementary drift indicators. The approach is evaluated across six malware families using fixed-interval and clustering-based windowing in family-vs-benign and family-vs-family settings, and compared against RIPPER and Transcendent baselines. Results show that fixed two-month windowing with feature-level Pearson correlation is the most reliable configuration, being the only one where all family pairs produce positive drift-accuracy correlations. The methods are complementary - no single approach dominates across all pairs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_22629
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Detecting Concept Drift in Evolving Malware Families Using Rule-Based Classifier Representations
Kalný, Tomáš
Jureček, Martin
Stamp, Mark
Cryptography and Security
Machine Learning
This work proposes a structural approach to concept drift detection in malware classification using decision tree rulesets. Classifiers are trained across temporal windows on the EMBER2024 dataset, and drift is quantified by comparing extracted rule representations using feature importance, prediction agreement, activation stability, and coverage metrics. These metrics are correlated with both accuracy degradation and data distribution shift as complementary drift indicators. The approach is evaluated across six malware families using fixed-interval and clustering-based windowing in family-vs-benign and family-vs-family settings, and compared against RIPPER and Transcendent baselines. Results show that fixed two-month windowing with feature-level Pearson correlation is the most reliable configuration, being the only one where all family pairs produce positive drift-accuracy correlations. The methods are complementary - no single approach dominates across all pairs.
title Detecting Concept Drift in Evolving Malware Families Using Rule-Based Classifier Representations
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2604.22629