Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gond, Bishwajit Prasad, Rajneekant, Kishore, Pushkar, Mohapatra, Durga Prasad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917284697604096
author Gond, Bishwajit Prasad
Rajneekant
Kishore, Pushkar
Mohapatra, Durga Prasad
author_facet Gond, Bishwajit Prasad
Rajneekant
Kishore, Pushkar
Mohapatra, Durga Prasad
contents This paper investigates the application of natural language processing (NLP)-based n-gram analysis and machine learning techniques to enhance malware classification. We explore how NLP can be used to extract and analyze textual features from malware samples through n-grams, contiguous string or API call sequences. This approach effectively captures distinctive linguistic patterns among malware and benign families, enabling finer-grained classification. We delve into n-gram size selection, feature representation, and classification algorithms. While evaluating our proposed method on real-world malware samples, we observe significantly improved accuracy compared to the traditional methods. By implementing our n-gram approach, we achieved an accuracy of 99.02% across various machine learning algorithms by using hybrid feature selection technique to address high dimensionality. Hybrid feature selection technique reduces the feature set to only 1.6% of the original features.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy
Gond, Bishwajit Prasad
Rajneekant
Kishore, Pushkar
Mohapatra, Durga Prasad
Cryptography and Security
Machine Learning
This paper investigates the application of natural language processing (NLP)-based n-gram analysis and machine learning techniques to enhance malware classification. We explore how NLP can be used to extract and analyze textual features from malware samples through n-grams, contiguous string or API call sequences. This approach effectively captures distinctive linguistic patterns among malware and benign families, enabling finer-grained classification. We delve into n-gram size selection, feature representation, and classification algorithms. While evaluating our proposed method on real-world malware samples, we observe significantly improved accuracy compared to the traditional methods. By implementing our n-gram approach, we achieved an accuracy of 99.02% across various machine learning algorithms by using hybrid feature selection technique to address high dimensionality. Hybrid feature selection technique reduces the feature set to only 1.6% of the original features.
title Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy
topic Cryptography and Security
Machine Learning
url https://arxiv.org/abs/2506.16224