Avoid Forgetting by Preserving Global Knowledge Gradients in Federated Learning with Non-IID Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chunduru, Abhijit, Morafah, Majid, Morafah, Mahdi, Chellapandi, Vishnu Pandi, Li, Ang
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909667779674112
author Chunduru, Abhijit
Morafah, Majid
Morafah, Mahdi
Chellapandi, Vishnu Pandi
Li, Ang
author_facet Chunduru, Abhijit
Morafah, Majid
Morafah, Mahdi
Chellapandi, Vishnu Pandi
Li, Ang
contents The inevitable presence of data heterogeneity has made federated learning very challenging. There are numerous methods to deal with this issue, such as local regularization, better model fusion techniques, and data sharing. Though effective, they lack a deep understanding of how data heterogeneity can affect the global decision boundary. In this paper, we bridge this gap by performing an experimental analysis of the learned decision boundary using a toy example. Our observations are surprising: (1) we find that the existing methods suffer from forgetting and clients forget the global decision boundary and only learn the perfect local one, and (2) this happens regardless of the initial weights, and clients forget the global decision boundary even starting from pre-trained optimal weights. In this paper, we present FedProj, a federated learning framework that robustly learns the global decision boundary and avoids its forgetting during local training. To achieve better ensemble knowledge fusion, we design a novel server-side ensemble knowledge transfer loss to further calibrate the learned global decision boundary. To alleviate the issue of learned global decision boundary forgetting, we further propose leveraging an episodic memory of average ensemble logits on a public unlabeled dataset to regulate the gradient updates at each step of local training. Experimental results demonstrate that FedProj outperforms state-of-the-art methods by a large margin.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Avoid Forgetting by Preserving Global Knowledge Gradients in Federated Learning with Non-IID Data
Chunduru, Abhijit
Morafah, Majid
Morafah, Mahdi
Chellapandi, Vishnu Pandi
Li, Ang
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
Performance
The inevitable presence of data heterogeneity has made federated learning very challenging. There are numerous methods to deal with this issue, such as local regularization, better model fusion techniques, and data sharing. Though effective, they lack a deep understanding of how data heterogeneity can affect the global decision boundary. In this paper, we bridge this gap by performing an experimental analysis of the learned decision boundary using a toy example. Our observations are surprising: (1) we find that the existing methods suffer from forgetting and clients forget the global decision boundary and only learn the perfect local one, and (2) this happens regardless of the initial weights, and clients forget the global decision boundary even starting from pre-trained optimal weights. In this paper, we present FedProj, a federated learning framework that robustly learns the global decision boundary and avoids its forgetting during local training. To achieve better ensemble knowledge fusion, we design a novel server-side ensemble knowledge transfer loss to further calibrate the learned global decision boundary. To alleviate the issue of learned global decision boundary forgetting, we further propose leveraging an episodic memory of average ensemble logits on a public unlabeled dataset to regulate the gradient updates at each step of local training. Experimental results demonstrate that FedProj outperforms state-of-the-art methods by a large margin.
title Avoid Forgetting by Preserving Global Knowledge Gradients in Federated Learning with Non-IID Data
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2505.20485