Scale-Robust Timely Asynchronous Decentralized Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mitra, Purbesh, Ulukus, Sennur
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913336727175168
author Mitra, Purbesh
Ulukus, Sennur
author_facet Mitra, Purbesh
Ulukus, Sennur
contents We consider an asynchronous decentralized learning system, which consists of a network of connected devices trying to learn a machine learning model without any centralized parameter server. The users in the network have their own local training data, which is used for learning across all the nodes in the network. The learning method consists of two processes, evolving simultaneously without any necessary synchronization. The first process is the model update, where the users update their local model via a fixed number of stochastic gradient descent steps. The second process is model mixing, where the users communicate with each other via randomized gossiping to exchange their models and average them to reach consensus. In this work, we investigate the staleness criteria for such a system, which is a sufficient condition for convergence of individual user models. We show that for network scaling, i.e., when the number of user devices $n$ is very large, if the gossip capacity of individual users scales as $Ω(\log n)$, we can guarantee the convergence of user models in finite time. Furthermore, we show that the bounded staleness can only be guaranteed by any distributed opportunistic scheme by $Ω(n)$ scaling.
format Preprint
id arxiv_https___arxiv_org_abs_2404_19749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scale-Robust Timely Asynchronous Decentralized Learning
Mitra, Purbesh
Ulukus, Sennur
Information Theory
Machine Learning
Multiagent Systems
Networking and Internet Architecture
Signal Processing
We consider an asynchronous decentralized learning system, which consists of a network of connected devices trying to learn a machine learning model without any centralized parameter server. The users in the network have their own local training data, which is used for learning across all the nodes in the network. The learning method consists of two processes, evolving simultaneously without any necessary synchronization. The first process is the model update, where the users update their local model via a fixed number of stochastic gradient descent steps. The second process is model mixing, where the users communicate with each other via randomized gossiping to exchange their models and average them to reach consensus. In this work, we investigate the staleness criteria for such a system, which is a sufficient condition for convergence of individual user models. We show that for network scaling, i.e., when the number of user devices $n$ is very large, if the gossip capacity of individual users scales as $Ω(\log n)$, we can guarantee the convergence of user models in finite time. Furthermore, we show that the bounded staleness can only be guaranteed by any distributed opportunistic scheme by $Ω(n)$ scaling.
title Scale-Robust Timely Asynchronous Decentralized Learning
topic Information Theory
Machine Learning
Multiagent Systems
Networking and Internet Architecture
Signal Processing
url https://arxiv.org/abs/2404.19749