Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yebo, Li, Jingguang, Tian, Chunlin, Tam, Kahou, Li, Li, Xu, Chengzhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918158865006592
author Wu, Yebo
Li, Jingguang
Tian, Chunlin
Tam, Kahou
Li, Li
Xu, Chengzhong
author_facet Wu, Yebo
Li, Jingguang
Tian, Chunlin
Tam, Kahou
Li, Li
Xu, Chengzhong
contents Federated Learning (FL) enables multiple clients to collaboratively train a shared model while preserving data privacy. However, the high memory demand during model training severely limits the deployment of FL on resource-constrained clients. To this end, we propose \our, a scalable and inclusive FL framework designed to overcome memory limitations through sequential block-wise training. The core idea of \our is to partition the global model into blocks and train them sequentially, thereby reducing training memory requirements. To mitigate information loss during block-wise training, \our introduces a Curriculum Mentor that crafts curriculum-aware training objectives for each block to steer their learning process. Moreover, \our incorporates a Training Harmonizer that designs a parameter co-adaptation training scheme to coordinate block updates, effectively breaking inter-block information isolation. Extensive experiments on both simulation and hardware testbeds demonstrate that \our significantly improves model performance by up to 84.2\%, reduces peak memory usage by up to 50.4\%, and accelerates training by up to 1.9$\times$.
format Preprint
id arxiv_https___arxiv_org_abs_2408_10826
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
Wu, Yebo
Li, Jingguang
Tian, Chunlin
Tam, Kahou
Li, Li
Xu, Chengzhong
Distributed, Parallel, and Cluster Computing
Federated Learning (FL) enables multiple clients to collaboratively train a shared model while preserving data privacy. However, the high memory demand during model training severely limits the deployment of FL on resource-constrained clients. To this end, we propose \our, a scalable and inclusive FL framework designed to overcome memory limitations through sequential block-wise training. The core idea of \our is to partition the global model into blocks and train them sequentially, thereby reducing training memory requirements. To mitigate information loss during block-wise training, \our introduces a Curriculum Mentor that crafts curriculum-aware training objectives for each block to steer their learning process. Moreover, \our incorporates a Training Harmonizer that designs a parameter co-adaptation training scheme to coordinate block updates, effectively breaking inter-block information isolation. Extensive experiments on both simulation and hardware testbeds demonstrate that \our significantly improves model performance by up to 84.2\%, reduces peak memory usage by up to 50.4\%, and accelerates training by up to 1.9$\times$.
title Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2408.10826