Revisiting Network Traffic Analysis: Compatible network flows for ML models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vitorino, João, Pinto, Daniela, Maia, Eva, Amorim, Ivone, Praça, Isabel
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915612044820480
author Vitorino, João
Pinto, Daniela
Maia, Eva
Amorim, Ivone
Praça, Isabel
author_facet Vitorino, João
Pinto, Daniela
Maia, Eva
Amorim, Ivone
Praça, Isabel
contents To ensure that Machine Learning (ML) models can perform a robust detection and classification of cyberattacks, it is essential to train them with high-quality datasets with relevant features. However, it can be difficult to accurately represent the complex traffic patterns of an attack, especially in Internet-of-Things (IoT) networks. This paper studies the impact that seemingly similar features created by different network traffic flow exporters can have on the generalization and robustness of ML models. In addition to the original CSV files of the Bot-IoT, IoT-23, and CICIoT23 datasets, the raw network packets of their PCAP files were analysed with the HERA tool, generating new labelled flows and extracting consistent features for new CSV versions. To assess the usefulness of these new flows for intrusion detection, they were compared with the original versions and were used to fine-tune multiple models. Overall, the results indicate that directly analysing and preprocessing PCAP files, instead of just using the commonly available CSV files, enables the computation of more relevant features to train bagging and gradient boosting decision tree ensembles. It is important to continue improving feature extraction and feature selection processes to make different datasets more compatible and enable a trustworthy evaluation and comparison of the ML models used in cybersecurity solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08345
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Network Traffic Analysis: Compatible network flows for ML models
Vitorino, João
Pinto, Daniela
Maia, Eva
Amorim, Ivone
Praça, Isabel
Cryptography and Security
Machine Learning
Networking and Internet Architecture
To ensure that Machine Learning (ML) models can perform a robust detection and classification of cyberattacks, it is essential to train them with high-quality datasets with relevant features. However, it can be difficult to accurately represent the complex traffic patterns of an attack, especially in Internet-of-Things (IoT) networks. This paper studies the impact that seemingly similar features created by different network traffic flow exporters can have on the generalization and robustness of ML models. In addition to the original CSV files of the Bot-IoT, IoT-23, and CICIoT23 datasets, the raw network packets of their PCAP files were analysed with the HERA tool, generating new labelled flows and extracting consistent features for new CSV versions. To assess the usefulness of these new flows for intrusion detection, they were compared with the original versions and were used to fine-tune multiple models. Overall, the results indicate that directly analysing and preprocessing PCAP files, instead of just using the commonly available CSV files, enables the computation of more relevant features to train bagging and gradient boosting decision tree ensembles. It is important to continue improving feature extraction and feature selection processes to make different datasets more compatible and enable a trustworthy evaluation and comparison of the ML models used in cybersecurity solutions.
title Revisiting Network Traffic Analysis: Compatible network flows for ML models
topic Cryptography and Security
Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2511.08345