Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917284697604096 |
|---|---|
| author | Gond, Bishwajit Prasad Rajneekant Kishore, Pushkar Mohapatra, Durga Prasad |
| author_facet | Gond, Bishwajit Prasad Rajneekant Kishore, Pushkar Mohapatra, Durga Prasad |
| contents | This paper investigates the application of natural language processing (NLP)-based n-gram analysis and machine learning techniques to enhance malware classification. We explore how NLP can be used to extract and analyze textual features from malware samples through n-grams, contiguous string or API call sequences. This approach effectively captures distinctive linguistic patterns among malware and benign families, enabling finer-grained classification. We delve into n-gram size selection, feature representation, and classification algorithms. While evaluating our proposed method on real-world malware samples, we observe significantly improved accuracy compared to the traditional methods. By implementing our n-gram approach, we achieved an accuracy of 99.02% across various machine learning algorithms by using hybrid feature selection technique to address high dimensionality. Hybrid feature selection technique reduces the feature set to only 1.6% of the original features. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_16224 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy Gond, Bishwajit Prasad Rajneekant Kishore, Pushkar Mohapatra, Durga Prasad Cryptography and Security Machine Learning This paper investigates the application of natural language processing (NLP)-based n-gram analysis and machine learning techniques to enhance malware classification. We explore how NLP can be used to extract and analyze textual features from malware samples through n-grams, contiguous string or API call sequences. This approach effectively captures distinctive linguistic patterns among malware and benign families, enabling finer-grained classification. We delve into n-gram size selection, feature representation, and classification algorithms. While evaluating our proposed method on real-world malware samples, we observe significantly improved accuracy compared to the traditional methods. By implementing our n-gram approach, we achieved an accuracy of 99.02% across various machine learning algorithms by using hybrid feature selection technique to address high dimensionality. Hybrid feature selection technique reduces the feature set to only 1.6% of the original features. |
| title | Malware Classification Leveraging NLP & Machine Learning for Enhanced Accuracy |
| topic | Cryptography and Security Machine Learning |
| url | https://arxiv.org/abs/2506.16224 |