Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Xinran, Javidi, Tara, Touri, Behrouz
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912939242422272
author Zheng, Xinran
Javidi, Tara
Touri, Behrouz
author_facet Zheng, Xinran
Javidi, Tara
Touri, Behrouz
contents We propose a general framework for distributed stochastic optimization under delayed gradient models. In this setting, $n$ local agents leverage their own data and computation to assist a central server in minimizing a global objective composed of agents' local cost functions. Each agent is allowed to transmit stochastic-potentially biased and delayed-estimates of its local gradient. While a prior work has advocated delay-adaptive step sizes for stochastic gradient descent (SGD) in the presence of delays, we demonstrate that a pre-chosen diminishing step size is sufficient and matches the performance of the adaptive scheme. Moreover, our analysis establishes that diminishing step sizes recover the optimal SGD rates for nonconvex and strongly convex objectives.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need
Zheng, Xinran
Javidi, Tara
Touri, Behrouz
Optimization and Control
Machine Learning
We propose a general framework for distributed stochastic optimization under delayed gradient models. In this setting, $n$ local agents leverage their own data and computation to assist a central server in minimizing a global objective composed of agents' local cost functions. Each agent is allowed to transmit stochastic-potentially biased and delayed-estimates of its local gradient. While a prior work has advocated delay-adaptive step sizes for stochastic gradient descent (SGD) in the presence of delays, we demonstrate that a pre-chosen diminishing step size is sufficient and matches the performance of the adaptive scheme. Moreover, our analysis establishes that diminishing step sizes recover the optimal SGD rates for nonconvex and strongly convex objectives.
title Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2603.02639