Federated Learning on Virtual Heterogeneous Data with Local-global Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Chun-Yin, Jin, Ruinan, Zhao, Can, Xu, Daguang, Li, Xiaoxiao
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910870563454976
author Huang, Chun-Yin
Jin, Ruinan
Zhao, Can
Xu, Daguang
Li, Xiaoxiao
author_facet Huang, Chun-Yin
Jin, Ruinan
Zhao, Can
Xu, Daguang
Li, Xiaoxiao
contents While Federated Learning (FL) is gaining popularity for training machine learning models in a decentralized fashion, numerous challenges persist, such as asynchronization, computational expenses, data heterogeneity, and gradient and membership privacy attacks. Lately, dataset distillation has emerged as a promising solution for addressing the aforementioned challenges by generating a compact synthetic dataset that preserves a model's training efficacy. However, we discover that using distilled local datasets can amplify the heterogeneity issue in FL. To address this, we propose Federated Learning on Virtual Heterogeneous Data with Local-Global Dataset Distillation (FedLGD), where we seamlessly integrate dataset distillation algorithms into FL pipeline and train FL using a smaller synthetic dataset (referred as virtual data). Specifically, to harmonize the domain shifts, we propose iterative distribution matching to inpaint global information to local virtual data and use federated gradient matching to distill global virtual data that serve as anchor points to rectify heterogeneous local training, without compromising data privacy. We experiment on both benchmark and real-world datasets that contain heterogeneous data from different sources, and further scale up to an FL scenario that contains a large number of clients with heterogeneous and class-imbalanced data. Our method outperforms state-of-the-art heterogeneous FL algorithms under various settings. Our code is available at https://github.com/ubc-tea/FedLGD.
format Preprint
id arxiv_https___arxiv_org_abs_2303_02278
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Federated Learning on Virtual Heterogeneous Data with Local-global Distillation
Huang, Chun-Yin
Jin, Ruinan
Zhao, Can
Xu, Daguang
Li, Xiaoxiao
Machine Learning
Artificial Intelligence
While Federated Learning (FL) is gaining popularity for training machine learning models in a decentralized fashion, numerous challenges persist, such as asynchronization, computational expenses, data heterogeneity, and gradient and membership privacy attacks. Lately, dataset distillation has emerged as a promising solution for addressing the aforementioned challenges by generating a compact synthetic dataset that preserves a model's training efficacy. However, we discover that using distilled local datasets can amplify the heterogeneity issue in FL. To address this, we propose Federated Learning on Virtual Heterogeneous Data with Local-Global Dataset Distillation (FedLGD), where we seamlessly integrate dataset distillation algorithms into FL pipeline and train FL using a smaller synthetic dataset (referred as virtual data). Specifically, to harmonize the domain shifts, we propose iterative distribution matching to inpaint global information to local virtual data and use federated gradient matching to distill global virtual data that serve as anchor points to rectify heterogeneous local training, without compromising data privacy. We experiment on both benchmark and real-world datasets that contain heterogeneous data from different sources, and further scale up to an FL scenario that contains a large number of clients with heterogeneous and class-imbalanced data. Our method outperforms state-of-the-art heterogeneous FL algorithms under various settings. Our code is available at https://github.com/ubc-tea/FedLGD.
title Federated Learning on Virtual Heterogeneous Data with Local-global Distillation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2303.02278