Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zheng, Ce, Ma, Shiyao, Zhang, Ke, Sun, Chen, Zhang, Wenqi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911386369523712
author Zheng, Ce
Ma, Shiyao
Zhang, Ke
Sun, Chen
Zhang, Wenqi
author_facet Zheng, Ce
Ma, Shiyao
Zhang, Ke
Sun, Chen
Zhang, Wenqi
contents Federated learning (FL) enables collaborative model training without sharing raw user data, but conventional simulations often rely on unrealistic data partitioning and current user selection methods ignore data correlation among users. To address these challenges, this paper proposes a metadatadriven FL framework. We first introduce a novel data partition model based on a homogeneous Poisson point process (HPPP), capturing both heterogeneity in data quantity and natural overlap among user datasets. Building on this model, we develop a clustering-based user selection strategy that leverages metadata, such as user location, to reduce data correlation and enhance label diversity across training rounds. Extensive experiments on FMNIST and CIFAR-10 demonstrate that the proposed framework improves model performance, stability, and convergence in non-IID scenarios, while maintaining comparable performance under IID settings. Furthermore, the method shows pronounced advantages when the number of selected users per round is small. These findings highlight the framework's potential for enhancing FL performance in realistic deployments and guiding future standardization.
format Preprint
id arxiv_https___arxiv_org_abs_2601_10013
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
Zheng, Ce
Ma, Shiyao
Zhang, Ke
Sun, Chen
Zhang, Wenqi
Signal Processing
Distributed, Parallel, and Cluster Computing
Federated learning (FL) enables collaborative model training without sharing raw user data, but conventional simulations often rely on unrealistic data partitioning and current user selection methods ignore data correlation among users. To address these challenges, this paper proposes a metadatadriven FL framework. We first introduce a novel data partition model based on a homogeneous Poisson point process (HPPP), capturing both heterogeneity in data quantity and natural overlap among user datasets. Building on this model, we develop a clustering-based user selection strategy that leverages metadata, such as user location, to reduce data correlation and enhance label diversity across training rounds. Extensive experiments on FMNIST and CIFAR-10 demonstrate that the proposed framework improves model performance, stability, and convergence in non-IID scenarios, while maintaining comparable performance under IID settings. Furthermore, the method shows pronounced advantages when the number of selected users per round is small. These findings highlight the framework's potential for enhancing FL performance in realistic deployments and guiding future standardization.
title Clustering-Based User Selection in Federated Learning: Metadata Exploitation for 3GPP Networks
topic Signal Processing
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.10013