StatAvg: Mitigating Data Heterogeneity in Federated Learning for Intrusion Detection Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bouzinis, Pavlos S., Radoglou-Grammatikis, Panagiotis, Makris, Ioannis, Lagkas, Thomas, Argyriou, Vasileios, Papadopoulos, Georgios Th., Sarigiannidis, Panagiotis, Karagiannidis, George K.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911884586778624
author Bouzinis, Pavlos S.
Radoglou-Grammatikis, Panagiotis
Makris, Ioannis
Lagkas, Thomas
Argyriou, Vasileios
Papadopoulos, Georgios Th.
Sarigiannidis, Panagiotis
Karagiannidis, George K.
author_facet Bouzinis, Pavlos S.
Radoglou-Grammatikis, Panagiotis
Makris, Ioannis
Lagkas, Thomas
Argyriou, Vasileios
Papadopoulos, Georgios Th.
Sarigiannidis, Panagiotis
Karagiannidis, George K.
contents Federated learning (FL) is a decentralized learning technique that enables participating devices to collaboratively build a shared Machine Leaning (ML) or Deep Learning (DL) model without revealing their raw data to a third party. Due to its privacy-preserving nature, FL has sparked widespread attention for building Intrusion Detection Systems (IDS) within the realm of cybersecurity. However, the data heterogeneity across participating domains and entities presents significant challenges for the reliable implementation of an FL-based IDS. In this paper, we propose an effective method called Statistical Averaging (StatAvg) to alleviate non-independently and identically (non-iid) distributed features across local clients' data in FL. In particular, StatAvg allows the FL clients to share their individual data statistics with the server, which then aggregates this information to produce global statistics. The latter are shared with the clients and used for universal data normalisation. It is worth mentioning that StatAvg can seamlessly integrate with any FL aggregation strategy, as it occurs before the actual FL training process. The proposed method is evaluated against baseline approaches using datasets for network and host Artificial Intelligence (AI)-powered IDS. The experimental results demonstrate the efficiency of StatAvg in mitigating non-iid feature distributions across the FL clients compared to the baseline methods.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13062
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle StatAvg: Mitigating Data Heterogeneity in Federated Learning for Intrusion Detection Systems
Bouzinis, Pavlos S.
Radoglou-Grammatikis, Panagiotis
Makris, Ioannis
Lagkas, Thomas
Argyriou, Vasileios
Papadopoulos, Georgios Th.
Sarigiannidis, Panagiotis
Karagiannidis, George K.
Cryptography and Security
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
Federated learning (FL) is a decentralized learning technique that enables participating devices to collaboratively build a shared Machine Leaning (ML) or Deep Learning (DL) model without revealing their raw data to a third party. Due to its privacy-preserving nature, FL has sparked widespread attention for building Intrusion Detection Systems (IDS) within the realm of cybersecurity. However, the data heterogeneity across participating domains and entities presents significant challenges for the reliable implementation of an FL-based IDS. In this paper, we propose an effective method called Statistical Averaging (StatAvg) to alleviate non-independently and identically (non-iid) distributed features across local clients' data in FL. In particular, StatAvg allows the FL clients to share their individual data statistics with the server, which then aggregates this information to produce global statistics. The latter are shared with the clients and used for universal data normalisation. It is worth mentioning that StatAvg can seamlessly integrate with any FL aggregation strategy, as it occurs before the actual FL training process. The proposed method is evaluated against baseline approaches using datasets for network and host Artificial Intelligence (AI)-powered IDS. The experimental results demonstrate the efficiency of StatAvg in mitigating non-iid feature distributions across the FL clients compared to the baseline methods.
title StatAvg: Mitigating Data Heterogeneity in Federated Learning for Intrusion Detection Systems
topic Cryptography and Security
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2405.13062