Stragglers-Aware Low-Latency Synchronous Federated Learning via Layer-Wise Model Updates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lang, Natalie, Cohen, Alejandro, Shlezinger, Nir
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910386335252480
author Lang, Natalie
Cohen, Alejandro
Shlezinger, Nir
author_facet Lang, Natalie
Cohen, Alejandro
Shlezinger, Nir
contents Synchronous federated learning (FL) is a popular paradigm for collaborative edge learning. It typically involves a set of heterogeneous devices locally training neural network (NN) models in parallel with periodic centralized aggregations. As some of the devices may have limited computational resources and varying availability, FL latency is highly sensitive to stragglers. Conventional approaches discard incomplete intra-model updates done by stragglers, alter the amount of local workload and architecture, or resort to asynchronous settings; which all affect the trained model performance under tight training latency constraints. In this work, we propose straggler-aware layer-wise federated learning (SALF) that leverages the optimization procedure of NNs via backpropagation to update the global model in a layer-wise fashion. SALF allows stragglers to synchronously convey partial gradients, having each layer of the global model be updated independently with a different contributing set of users. We provide a theoretical analysis, establishing convergence guarantees for the global model under mild assumptions on the distribution of the participating devices, revealing that SALF converges at the same asymptotic rate as FL with no timing limitations. This insight is matched with empirical observations, demonstrating the performance gains of SALF compared to alternative mechanisms mitigating the device heterogeneity gap in FL.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18375
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Stragglers-Aware Low-Latency Synchronous Federated Learning via Layer-Wise Model Updates
Lang, Natalie
Cohen, Alejandro
Shlezinger, Nir
Machine Learning
Signal Processing
Synchronous federated learning (FL) is a popular paradigm for collaborative edge learning. It typically involves a set of heterogeneous devices locally training neural network (NN) models in parallel with periodic centralized aggregations. As some of the devices may have limited computational resources and varying availability, FL latency is highly sensitive to stragglers. Conventional approaches discard incomplete intra-model updates done by stragglers, alter the amount of local workload and architecture, or resort to asynchronous settings; which all affect the trained model performance under tight training latency constraints. In this work, we propose straggler-aware layer-wise federated learning (SALF) that leverages the optimization procedure of NNs via backpropagation to update the global model in a layer-wise fashion. SALF allows stragglers to synchronously convey partial gradients, having each layer of the global model be updated independently with a different contributing set of users. We provide a theoretical analysis, establishing convergence guarantees for the global model under mild assumptions on the distribution of the participating devices, revealing that SALF converges at the same asymptotic rate as FL with no timing limitations. This insight is matched with empirical observations, demonstrating the performance gains of SALF compared to alternative mechanisms mitigating the device heterogeneity gap in FL.
title Stragglers-Aware Low-Latency Synchronous Federated Learning via Layer-Wise Model Updates
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2403.18375