Saved in:
Bibliographic Details
Main Authors: Fu, Annan, Pei, Hao, Tanha, Maryam
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.23025
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918466715385856
author Fu, Annan
Pei, Hao
Tanha, Maryam
author_facet Fu, Annan
Pei, Hao
Tanha, Maryam
contents Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening real-world robustness. We address this by constructing a time-stamped dataset of benign and malicious Android apps and introducing a timestamp-verification procedure to ensure temporal accuracy. We then propose a detection framework that uses Bootstrap Your Own Latent (BYOL) for self-supervised pre-training to learn obfuscation-resilient representations, followed by supervised classification. Under time-aware evaluation, the method attains 98% accuracy and 89% F1. We further characterize malware behavior by analyzing true positives and false negatives using VirusTotal and the MITRE ATT&CK framework. To support reproducibility and further innovation, we release our dataset and source code.
format Preprint
id arxiv_https___arxiv_org_abs_2604_23025
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Self-Supervised Learning for Android Malware Detection on a Time-Stamped Dataset
Fu, Annan
Pei, Hao
Tanha, Maryam
Cryptography and Security
Machine Learning
Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening real-world robustness. We address this by constructing a time-stamped dataset of benign and malicious Android apps and introducing a timestamp-verification procedure to ensure temporal accuracy. We then propose a detection framework that uses Bootstrap Your Own Latent (BYOL) for self-supervised pre-training to learn obfuscation-resilient representations, followed by supervised classification. Under time-aware evaluation, the method attains 98% accuracy and 89% F1. We further characterize malware behavior by analyzing true positives and false negatives using VirusTotal and the MITRE ATT&CK framework. To support reproducibility and further innovation, we release our dataset and source code.
title Self-Supervised Learning for Android Malware Detection on a Time-Stamped Dataset
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2604.23025