AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ajanthan, Thalaiyasingam, Ramasinghe, Sameera, Avraham, Gil, Dolatabadi, Hadi Mohaghegh, Koneputugodage, Chamin P Hewa, Shevchenko, Violetta, Zuo, Yan, Long, Alexander
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910005448409088
author Ajanthan, Thalaiyasingam
Ramasinghe, Sameera
Avraham, Gil
Dolatabadi, Hadi Mohaghegh
Koneputugodage, Chamin P Hewa
Shevchenko, Violetta
Zuo, Yan
Long, Alexander
author_facet Ajanthan, Thalaiyasingam
Ramasinghe, Sameera
Avraham, Gil
Dolatabadi, Hadi Mohaghegh
Koneputugodage, Chamin P Hewa
Shevchenko, Violetta
Zuo, Yan
Long, Alexander
contents Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing asynchronous updates across both parallelism axes, relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an asynchronous sparse averaging method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models (up to \em 1B parameters) demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22442
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
Ajanthan, Thalaiyasingam
Ramasinghe, Sameera
Avraham, Gil
Dolatabadi, Hadi Mohaghegh
Koneputugodage, Chamin P Hewa
Shevchenko, Violetta
Zuo, Yan
Long, Alexander
Machine Learning
Distributed, Parallel, and Cluster Computing
Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing asynchronous updates across both parallelism axes, relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an asynchronous sparse averaging method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models (up to \em 1B parameters) demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead.
title AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
topic Machine Learning
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.22442