Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xiong, Guojun, Yan, Gang, Wang, Shiqiang, Li, Jian
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913421436387328
author Xiong, Guojun
Yan, Gang
Wang, Shiqiang
Li, Jian
author_facet Xiong, Guojun
Yan, Gang
Wang, Shiqiang
Li, Jian
contents With the increasing demand for large-scale training of machine learning models, fully decentralized optimization methods have recently been advocated as alternatives to the popular parameter server framework. In this paradigm, each worker maintains a local estimate of the optimal parameter vector, and iteratively updates it by waiting and averaging all estimates obtained from its neighbors, and then corrects it on the basis of its local dataset. However, the synchronization phase is sensitive to stragglers. An efficient way to mitigate this effect is to consider asynchronous updates, where each worker computes stochastic gradients and communicates with other workers at its own pace. Unfortunately, fully asynchronous updates suffer from staleness of stragglers' parameters. To address these limitations, we propose a fully decentralized algorithm DSGD-AAU with adaptive asynchronous updates via adaptively determining the number of neighbor workers for each worker to communicate with. We show that DSGD-AAU achieves a linear speedup for convergence and demonstrate its effectiveness via extensive experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2306_06559
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
Xiong, Guojun
Yan, Gang
Wang, Shiqiang
Li, Jian
Machine Learning
Distributed, Parallel, and Cluster Computing
With the increasing demand for large-scale training of machine learning models, fully decentralized optimization methods have recently been advocated as alternatives to the popular parameter server framework. In this paradigm, each worker maintains a local estimate of the optimal parameter vector, and iteratively updates it by waiting and averaging all estimates obtained from its neighbors, and then corrects it on the basis of its local dataset. However, the synchronization phase is sensitive to stragglers. An efficient way to mitigate this effect is to consider asynchronous updates, where each worker computes stochastic gradients and communicates with other workers at its own pace. Unfortunately, fully asynchronous updates suffer from staleness of stragglers' parameters. To address these limitations, we propose a fully decentralized algorithm DSGD-AAU with adaptive asynchronous updates via adaptively determining the number of neighbor workers for each worker to communicate with. We show that DSGD-AAU achieves a linear speedup for convergence and demonstrate its effectiveness via extensive experiments.
title Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2306.06559