Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.23025 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918466715385856 |
|---|---|
| author | Fu, Annan Pei, Hao Tanha, Maryam |
| author_facet | Fu, Annan Pei, Hao Tanha, Maryam |
| contents | Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening real-world robustness. We address this by constructing a time-stamped dataset of benign and malicious Android apps and introducing a timestamp-verification procedure to ensure temporal accuracy. We then propose a detection framework that uses Bootstrap Your Own Latent (BYOL) for self-supervised pre-training to learn obfuscation-resilient representations, followed by supervised classification. Under time-aware evaluation, the method attains 98% accuracy and 89% F1. We further characterize malware behavior by analyzing true positives and false negatives using VirusTotal and the MITRE ATT&CK framework. To support reproducibility and further innovation, we release our dataset and source code. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_23025 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Self-Supervised Learning for Android Malware Detection on a Time-Stamped Dataset Fu, Annan Pei, Hao Tanha, Maryam Cryptography and Security Machine Learning Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening real-world robustness. We address this by constructing a time-stamped dataset of benign and malicious Android apps and introducing a timestamp-verification procedure to ensure temporal accuracy. We then propose a detection framework that uses Bootstrap Your Own Latent (BYOL) for self-supervised pre-training to learn obfuscation-resilient representations, followed by supervised classification. Under time-aware evaluation, the method attains 98% accuracy and 89% F1. We further characterize malware behavior by analyzing true positives and false negatives using VirusTotal and the MITRE ATT&CK framework. To support reproducibility and further innovation, we release our dataset and source code. |
| title | Self-Supervised Learning for Android Malware Detection on a Time-Stamped Dataset |
| topic | Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2604.23025 |