BEACON: Behavioral Malware Classification with Large Language Model Embeddings and Deep Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Perera, Wadduwage Shanika, Jiang, Haodi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909794647932928
author Perera, Wadduwage Shanika
Jiang, Haodi
author_facet Perera, Wadduwage Shanika
Jiang, Haodi
contents Malware is becoming increasingly complex and widespread, making it essential to develop more effective and timely detection methods. Traditional static analysis often fails to defend against modern threats that employ code obfuscation, polymorphism, and other evasion techniques. In contrast, behavioral malware detection, which monitors runtime activities, provides a more reliable and context-aware solution. In this work, we propose BEACON, a novel deep learning framework that leverages large language models (LLMs) to generate dense, contextual embeddings from raw sandbox-generated behavior reports. These embeddings capture semantic and structural patterns of each sample and are processed by a one-dimensional convolutional neural network (1D CNN) for multi-class malware classification. Evaluated on the Avast-CTU Public CAPE Dataset, our framework consistently outperforms existing methods, highlighting the effectiveness of LLM-based behavioral embeddings and the overall design of BEACON for robust malware classification.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14519
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BEACON: Behavioral Malware Classification with Large Language Model Embeddings and Deep Learning
Perera, Wadduwage Shanika
Jiang, Haodi
Machine Learning
Artificial Intelligence
Cryptography and Security
Malware is becoming increasingly complex and widespread, making it essential to develop more effective and timely detection methods. Traditional static analysis often fails to defend against modern threats that employ code obfuscation, polymorphism, and other evasion techniques. In contrast, behavioral malware detection, which monitors runtime activities, provides a more reliable and context-aware solution. In this work, we propose BEACON, a novel deep learning framework that leverages large language models (LLMs) to generate dense, contextual embeddings from raw sandbox-generated behavior reports. These embeddings capture semantic and structural patterns of each sample and are processed by a one-dimensional convolutional neural network (1D CNN) for multi-class malware classification. Evaluated on the Avast-CTU Public CAPE Dataset, our framework consistently outperforms existing methods, highlighting the effectiveness of LLM-based behavioral embeddings and the overall design of BEACON for robust malware classification.
title BEACON: Behavioral Malware Classification with Large Language Model Embeddings and Deep Learning
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2509.14519