AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910005448409088 |
|---|---|
| author | Ajanthan, Thalaiyasingam Ramasinghe, Sameera Avraham, Gil Dolatabadi, Hadi Mohaghegh Koneputugodage, Chamin P Hewa Shevchenko, Violetta Zuo, Yan Long, Alexander |
| author_facet | Ajanthan, Thalaiyasingam Ramasinghe, Sameera Avraham, Gil Dolatabadi, Hadi Mohaghegh Koneputugodage, Chamin P Hewa Shevchenko, Violetta Zuo, Yan Long, Alexander |
| contents | Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing asynchronous updates across both parallelism axes, relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an asynchronous sparse averaging method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models (up to \em 1B parameters) demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_22442 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism Ajanthan, Thalaiyasingam Ramasinghe, Sameera Avraham, Gil Dolatabadi, Hadi Mohaghegh Koneputugodage, Chamin P Hewa Shevchenko, Violetta Zuo, Yan Long, Alexander Machine Learning Distributed, Parallel, and Cluster Computing Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing asynchronous updates across both parallelism axes, relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an asynchronous sparse averaging method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models (up to \em 1B parameters) demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead. |
| title | AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism |
| topic | Machine Learning Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2601.22442 |