Comparative Analysis of Hash-based Malware Clustering via K-Means

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thein, Aink Acrie Soe, Pitropakis, Nikolaos, Papadopoulos, Pavlos, Grierson, Sam, Jan, Sana Ullah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911311242199040
author Thein, Aink Acrie Soe
Pitropakis, Nikolaos
Papadopoulos, Pavlos
Grierson, Sam
Jan, Sana Ullah
author_facet Thein, Aink Acrie Soe
Pitropakis, Nikolaos
Papadopoulos, Pavlos
Grierson, Sam
Jan, Sana Ullah
contents With the adoption of multiple digital devices in everyday life, the cyber-attack surface has increased. Adversaries are continuously exploring new avenues to exploit them and deploy malware. On the other hand, detection approaches typically employ hashing-based algorithms such as SSDeep, TLSH, and IMPHash to capture structural and behavioural similarities among binaries. This work focuses on the analysis and evaluation of these techniques for clustering malware samples using the K-means algorithm. More specifically, we experimented with established malware families and traits and found that TLSH and IMPHash produce more distinct, semantically meaningful clusters, whereas SSDeep is more efficient for broader classification tasks. The findings of this work can guide the development of more robust threat-detection mechanisms and adaptive security mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2512_09539
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparative Analysis of Hash-based Malware Clustering via K-Means
Thein, Aink Acrie Soe
Pitropakis, Nikolaos
Papadopoulos, Pavlos
Grierson, Sam
Jan, Sana Ullah
Cryptography and Security
Machine Learning
With the adoption of multiple digital devices in everyday life, the cyber-attack surface has increased. Adversaries are continuously exploring new avenues to exploit them and deploy malware. On the other hand, detection approaches typically employ hashing-based algorithms such as SSDeep, TLSH, and IMPHash to capture structural and behavioural similarities among binaries. This work focuses on the analysis and evaluation of these techniques for clustering malware samples using the K-means algorithm. More specifically, we experimented with established malware families and traits and found that TLSH and IMPHash produce more distinct, semantically meaningful clusters, whereas SSDeep is more efficient for broader classification tasks. The findings of this work can guide the development of more robust threat-detection mechanisms and adaptive security mechanisms.
title Comparative Analysis of Hash-based Malware Clustering via K-Means
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2512.09539