Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mohammed, Riyazuddin, Zhang, Lan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913101333397504
author Mohammed, Riyazuddin
Zhang, Lan
author_facet Mohammed, Riyazuddin
Zhang, Lan
contents Modern cybersecurity relies heavily on static machine-learning-based malware classifiers. However, transformations such as packing and other non-semantic modifications applied to executable files limit their reliability. Malware classifiers often learn these unnecessary artifacts rather than the true binary behavior because of the high association between maliciousness and packing. Moreover, these malware classifiers are black boxes, making it difficult to understand what they learn. To address this issue, we proposed a two-part framework using the post-hoc interpretability XAI tool TRUSTEE, followed by a manual analysis of the top features. We conducted several controlled experiments by varying the dataset composition ratios to understand their impact on the results. The top-ranked features across all experiments, identified by TRUSTEE, were predominantly packing artifacts, portable executable(PE) metadata, and n-grams at the string level, rather than malicious semantics. These results suggest that these malware classifiers are highly sensitive to dataset composition and can misinterpret packing as malicious behavior. Our proposed framework allows for the reproducible diagnosis of such biases and forms a guideline for building more robust and semantically meaningful malware detection models
format Preprint
id arxiv_https___arxiv_org_abs_2605_07034
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
Mohammed, Riyazuddin
Zhang, Lan
Cryptography and Security
Machine Learning
Modern cybersecurity relies heavily on static machine-learning-based malware classifiers. However, transformations such as packing and other non-semantic modifications applied to executable files limit their reliability. Malware classifiers often learn these unnecessary artifacts rather than the true binary behavior because of the high association between maliciousness and packing. Moreover, these malware classifiers are black boxes, making it difficult to understand what they learn. To address this issue, we proposed a two-part framework using the post-hoc interpretability XAI tool TRUSTEE, followed by a manual analysis of the top features. We conducted several controlled experiments by varying the dataset composition ratios to understand their impact on the results. The top-ranked features across all experiments, identified by TRUSTEE, were predominantly packing artifacts, portable executable(PE) metadata, and n-grams at the string level, rather than malicious semantics. These results suggest that these malware classifiers are highly sensitive to dataset composition and can misinterpret packing as malicious behavior. Our proposed framework allows for the reproducible diagnosis of such biases and forms a guideline for building more robust and semantically meaningful malware detection models
title Beyond the Wrapper: Identifying Artifact Reliance in Static Malware Classifiers using TRUSTEE
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2605.07034