FedClust: Tackling Data Heterogeneity in Federated Learning through Weight-Driven Client Clustering

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Islam, Md Sirajul, Javaherian, Simin, Xu, Fei, Yuan, Xu, Chen, Li, Tzeng, Nian-Feng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909249452376064
author Islam, Md Sirajul
Javaherian, Simin
Xu, Fei
Yuan, Xu
Chen, Li
Tzeng, Nian-Feng
author_facet Islam, Md Sirajul
Javaherian, Simin
Xu, Fei
Yuan, Xu
Chen, Li
Tzeng, Nian-Feng
contents Federated learning (FL) is an emerging distributed machine learning paradigm that enables collaborative training of machine learning models over decentralized devices without exposing their local data. One of the major challenges in FL is the presence of uneven data distributions across client devices, violating the well-known assumption of independent-and-identically-distributed (IID) training samples in conventional machine learning. To address the performance degradation issue incurred by such data heterogeneity, clustered federated learning (CFL) shows its promise by grouping clients into separate learning clusters based on the similarity of their local data distributions. However, state-of-the-art CFL approaches require a large number of communication rounds to learn the distribution similarities during training until the formation of clusters is stabilized. Moreover, some of these algorithms heavily rely on a predefined number of clusters, thus limiting their flexibility and adaptability. In this paper, we propose {\em FedClust}, a novel approach for CFL that leverages the correlation between local model weights and the data distribution of clients. {\em FedClust} groups clients into clusters in a one-shot manner by measuring the similarity degrees among clients based on the strategically selected partial weights of locally trained models. We conduct extensive experiments on four benchmark datasets with different non-IID data settings. Experimental results demonstrate that {\em FedClust} achieves higher model accuracy up to $\sim$45\% as well as faster convergence with a significantly reduced communication cost up to 2.7$\times$ compared to its state-of-the-art counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2407_07124
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FedClust: Tackling Data Heterogeneity in Federated Learning through Weight-Driven Client Clustering
Islam, Md Sirajul
Javaherian, Simin
Xu, Fei
Yuan, Xu
Chen, Li
Tzeng, Nian-Feng
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
Federated learning (FL) is an emerging distributed machine learning paradigm that enables collaborative training of machine learning models over decentralized devices without exposing their local data. One of the major challenges in FL is the presence of uneven data distributions across client devices, violating the well-known assumption of independent-and-identically-distributed (IID) training samples in conventional machine learning. To address the performance degradation issue incurred by such data heterogeneity, clustered federated learning (CFL) shows its promise by grouping clients into separate learning clusters based on the similarity of their local data distributions. However, state-of-the-art CFL approaches require a large number of communication rounds to learn the distribution similarities during training until the formation of clusters is stabilized. Moreover, some of these algorithms heavily rely on a predefined number of clusters, thus limiting their flexibility and adaptability. In this paper, we propose {\em FedClust}, a novel approach for CFL that leverages the correlation between local model weights and the data distribution of clients. {\em FedClust} groups clients into clusters in a one-shot manner by measuring the similarity degrees among clients based on the strategically selected partial weights of locally trained models. We conduct extensive experiments on four benchmark datasets with different non-IID data settings. Experimental results demonstrate that {\em FedClust} achieves higher model accuracy up to $\sim$45\% as well as faster convergence with a significantly reduced communication cost up to 2.7$\times$ compared to its state-of-the-art counterparts.
title FedClust: Tackling Data Heterogeneity in Federated Learning through Weight-Driven Client Clustering
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.07124