Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meinhardt, Georg, Yi, Kai, Condat, Laurent, Richtárik, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909515742445568
author Meinhardt, Georg
Yi, Kai
Condat, Laurent
Richtárik, Peter
author_facet Meinhardt, Georg
Yi, Kai
Condat, Laurent
Richtárik, Peter
contents In Federated Learning (FL), both client resource constraints and communication costs pose major problems for training large models. In the centralized setting, sparse training addresses resource constraints, while in the distributed setting, local training addresses communication costs. Recent work has shown that local training provably improves communication complexity through acceleration. In this work we show that in FL, naive integration of sparse training and acceleration fails, and we provide theoretical and empirical explanations of this phenomenon. We introduce Sparse-ProxSkip, addressing the issue and implementing the efficient technique of Straight-Through Estimator pruning into sparse training. We demonstrate the performance of Sparse-ProxSkip in extensive experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20623
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning
Meinhardt, Georg
Yi, Kai
Condat, Laurent
Richtárik, Peter
Machine Learning
Optimization and Control
In Federated Learning (FL), both client resource constraints and communication costs pose major problems for training large models. In the centralized setting, sparse training addresses resource constraints, while in the distributed setting, local training addresses communication costs. Recent work has shown that local training provably improves communication complexity through acceleration. In this work we show that in FL, naive integration of sparse training and acceleration fails, and we provide theoretical and empirical explanations of this phenomenon. We introduce Sparse-ProxSkip, addressing the issue and implementing the efficient technique of Straight-Through Estimator pruning into sparse training. We demonstrate the performance of Sparse-ProxSkip in extensive experiments.
title Sparse-ProxSkip: Accelerated Sparse-to-Sparse Training in Federated Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2405.20623