Generic Multi-modal Representation Learning for Network Traffic Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gioacchini, Luca, Drago, Idilio, Mellia, Marco, Houidi, Zied Ben, Rossi, Dario
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914783978061824
author Gioacchini, Luca
Drago, Idilio
Mellia, Marco
Houidi, Zied Ben
Rossi, Dario
author_facet Gioacchini, Luca
Drago, Idilio
Mellia, Marco
Houidi, Zied Ben
Rossi, Dario
contents Network traffic analysis is fundamental for network management, troubleshooting, and security. Tasks such as traffic classification, anomaly detection, and novelty discovery are fundamental for extracting operational information from network data and measurements. We witness the shift from deep packet inspection and basic machine learning to Deep Learning (DL) approaches where researchers define and test a custom DL architecture designed for each specific problem. We here advocate the need for a general DL architecture flexible enough to solve different traffic analysis tasks. We test this idea by proposing a DL architecture based on generic data adaptation modules, followed by an integration module that summarises the extracted information into a compact and rich intermediate representation (i.e. embeddings). The result is a flexible Multi-modal Autoencoder (MAE) pipeline that can solve different use cases. We demonstrate the architecture with traffic classification (TC) tasks since they allow us to quantitatively compare results with state-of-the-art solutions. However, we argue that the MAE architecture is generic and can be used to learn representations useful in multiple scenarios. On TC, the MAE performs on par or better than alternatives while avoiding cumbersome feature engineering, thus streamlining the adoption of DL solutions for traffic analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2405_02649
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generic Multi-modal Representation Learning for Network Traffic Analysis
Gioacchini, Luca
Drago, Idilio
Mellia, Marco
Houidi, Zied Ben
Rossi, Dario
Machine Learning
Artificial Intelligence
Network traffic analysis is fundamental for network management, troubleshooting, and security. Tasks such as traffic classification, anomaly detection, and novelty discovery are fundamental for extracting operational information from network data and measurements. We witness the shift from deep packet inspection and basic machine learning to Deep Learning (DL) approaches where researchers define and test a custom DL architecture designed for each specific problem. We here advocate the need for a general DL architecture flexible enough to solve different traffic analysis tasks. We test this idea by proposing a DL architecture based on generic data adaptation modules, followed by an integration module that summarises the extracted information into a compact and rich intermediate representation (i.e. embeddings). The result is a flexible Multi-modal Autoencoder (MAE) pipeline that can solve different use cases. We demonstrate the architecture with traffic classification (TC) tasks since they allow us to quantitatively compare results with state-of-the-art solutions. However, we argue that the MAE architecture is generic and can be used to learn representations useful in multiple scenarios. On TC, the MAE performs on par or better than alternatives while avoiding cumbersome feature engineering, thus streamlining the adoption of DL solutions for traffic analysis.
title Generic Multi-modal Representation Learning for Network Traffic Analysis
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.02649