Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Subedar, Noah, Kim, Taeui, Venkataramalingam, Saathwick
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910916265639936
author Subedar, Noah
Kim, Taeui
Venkataramalingam, Saathwick
author_facet Subedar, Noah
Kim, Taeui
Venkataramalingam, Saathwick
contents This paper presents an underlying framework for both automating and accelerating malware classification, more specifically, mapping malicious executables to known Advanced Persistent Threat (APT) groups. The main feature of this analysis is the assembly-level instructions present in executables which are also known as opcodes. The collection of such opcodes on many malicious samples is a lengthy process; hence, open-source reverse engineering tools are used in tandem with scripts that leverage parallel computing to analyze multiple files at once. Traditional and deep learning models are applied to create models capable of classifying malware samples. One-gram and two-gram datasets are constructed and used to train models such as SVM, KNN, and Decision Tree; however, they struggle to provide adequate results without relying on metadata to support n-gram sequences. The computational limitations of such models are overcome with convolutional neural networks (CNNs) and heavily accelerated using graphical compute unit (GPU) resources.
format Preprint
id arxiv_https___arxiv_org_abs_2504_15497
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
Subedar, Noah
Kim, Taeui
Venkataramalingam, Saathwick
Cryptography and Security
Artificial Intelligence
I.2.0; I.2.6; K.6.5
This paper presents an underlying framework for both automating and accelerating malware classification, more specifically, mapping malicious executables to known Advanced Persistent Threat (APT) groups. The main feature of this analysis is the assembly-level instructions present in executables which are also known as opcodes. The collection of such opcodes on many malicious samples is a lengthy process; hence, open-source reverse engineering tools are used in tandem with scripts that leverage parallel computing to analyze multiple files at once. Traditional and deep learning models are applied to create models capable of classifying malware samples. One-gram and two-gram datasets are constructed and used to train models such as SVM, KNN, and Decision Tree; however, they struggle to provide adequate results without relying on metadata to support n-gram sequences. The computational limitations of such models are overcome with convolutional neural networks (CNNs) and heavily accelerated using graphical compute unit (GPU) resources.
title Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning
topic Cryptography and Security
Artificial Intelligence
I.2.0; I.2.6; K.6.5
url https://arxiv.org/abs/2504.15497