Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Antari, Ahmad, Abo-Aisheh, Yazan, Shamasneh, Jehad, Ashqar, Huthaifa I.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912257679556608
author Antari, Ahmad
Abo-Aisheh, Yazan
Shamasneh, Jehad
Ashqar, Huthaifa I.
author_facet Antari, Ahmad
Abo-Aisheh, Yazan
Shamasneh, Jehad
Ashqar, Huthaifa I.
contents This study uses various models to address network traffic classification, categorizing traffic into web, browsing, IPSec, backup, and email. We collected a comprehensive dataset from Arbor Edge Defender (AED) devices, comprising of 30,959 observations and 19 features. Multiple models were evaluated, including Naive Bayes, Decision Tree, Random Forest, Gradient Boosting, XGBoost, Deep Neural Networks (DNN), Transformer, and two Large Language Models (LLMs) including GPT-4o and Gemini with zero- and few-shot learning. Transformer and XGBoost showed the best performance, achieving the highest accuracy of 98.95 and 97.56%, respectively. GPT-4o and Gemini showed promising results with few-shot learning, improving accuracy significantly from initial zero-shot performance. While Gemini Few-Shot and GPT-4o Few-Shot performed well in categories like Web and Email, misclassifications occurred in more complex categories like IPSec and Backup. The study highlights the importance of model selection, fine-tuning, and the balance between training data size and model complexity for achieving reliable classification results.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02141
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
Antari, Ahmad
Abo-Aisheh, Yazan
Shamasneh, Jehad
Ashqar, Huthaifa I.
Machine Learning
Computation and Language
Cryptography and Security
This study uses various models to address network traffic classification, categorizing traffic into web, browsing, IPSec, backup, and email. We collected a comprehensive dataset from Arbor Edge Defender (AED) devices, comprising of 30,959 observations and 19 features. Multiple models were evaluated, including Naive Bayes, Decision Tree, Random Forest, Gradient Boosting, XGBoost, Deep Neural Networks (DNN), Transformer, and two Large Language Models (LLMs) including GPT-4o and Gemini with zero- and few-shot learning. Transformer and XGBoost showed the best performance, achieving the highest accuracy of 98.95 and 97.56%, respectively. GPT-4o and Gemini showed promising results with few-shot learning, improving accuracy significantly from initial zero-shot performance. While Gemini Few-Shot and GPT-4o Few-Shot performed well in categories like Web and Email, misclassifications occurred in more complex categories like IPSec and Backup. The study highlights the importance of model selection, fine-tuning, and the balance between training data size and model complexity for achieving reliable classification results.
title Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
topic Machine Learning
Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2503.02141