Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Zhongwei, Wu, Tong, Chen, Zhiyong, Qian, Liang, Xu, Yin, Tao, Meixia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929658072662016
author Wang, Zhongwei
Wu, Tong
Chen, Zhiyong
Qian, Liang
Xu, Yin
Tao, Meixia
author_facet Wang, Zhongwei
Wu, Tong
Chen, Zhiyong
Qian, Liang
Xu, Yin
Tao, Meixia
contents Federated semi-supervised learning (FSSL) is primarily challenged by two factors: the scarcity of labeled data across clients and the non-independent and identically distribution (non-IID) nature of data among clients. In this paper, we propose a novel approach, diffusion model-based data synthesis aided FSSL (DDSA-FSSL), which utilizes a diffusion model (DM) to generate synthetic data, bridging the gap between heterogeneous local data distributions and the global data distribution. In DDSA-FSSL, clients address the challenge of the scarcity of labeled data by employing a federated learning-trained classifier to perform pseudo labeling for unlabeled data. The DM is then collaboratively trained using both labeled and precision-optimized pseudo-labeled data, enabling clients to generate synthetic samples for classes that are absent in their labeled datasets. This process allows clients to generate more comprehensive synthetic datasets aligned with the global distribution. Extensive experiments conducted on multiple datasets and varying non-IID distributions demonstrate the effectiveness of DDSA-FSSL, e.g., it improves accuracy from 38.46% to 52.14% on CIFAR-10 datasets with 10% labeled data.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02219
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
Wang, Zhongwei
Wu, Tong
Chen, Zhiyong
Qian, Liang
Xu, Yin
Tao, Meixia
Machine Learning
Artificial Intelligence
Information Theory
Federated semi-supervised learning (FSSL) is primarily challenged by two factors: the scarcity of labeled data across clients and the non-independent and identically distribution (non-IID) nature of data among clients. In this paper, we propose a novel approach, diffusion model-based data synthesis aided FSSL (DDSA-FSSL), which utilizes a diffusion model (DM) to generate synthetic data, bridging the gap between heterogeneous local data distributions and the global data distribution. In DDSA-FSSL, clients address the challenge of the scarcity of labeled data by employing a federated learning-trained classifier to perform pseudo labeling for unlabeled data. The DM is then collaboratively trained using both labeled and precision-optimized pseudo-labeled data, enabling clients to generate synthetic samples for classes that are absent in their labeled datasets. This process allows clients to generate more comprehensive synthetic datasets aligned with the global distribution. Extensive experiments conducted on multiple datasets and varying non-IID distributions demonstrate the effectiveness of DDSA-FSSL, e.g., it improves accuracy from 38.46% to 52.14% on CIFAR-10 datasets with 10% labeled data.
title Diffusion Model-Based Data Synthesis Aided Federated Semi-Supervised Learning
topic Machine Learning
Artificial Intelligence
Information Theory
url https://arxiv.org/abs/2501.02219