Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ramasinghe, Sameera, Ajanthan, Thalaiyasingam, Avraham, Gil, Zuo, Yan, Long, Alexander
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918042823294976
author Ramasinghe, Sameera
Ajanthan, Thalaiyasingam
Avraham, Gil
Zuo, Yan
Long, Alexander
author_facet Ramasinghe, Sameera
Ajanthan, Thalaiyasingam
Avraham, Gil
Zuo, Yan
Long, Alexander
contents Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks. While existing compression techniques are effective in data-parallel, they do not extend to model parallelism. Unlike data-parallel training, where weight gradients are exchanged, model-parallel requires compressing activations and activation gradients as they propagate through layers, accumulating compression errors. We propose a novel compression algorithm that compresses both forward and backward passes, enabling up to 99% compression with no convergence degradation with negligible memory/compute overhead. By leveraging a recursive structure in transformer networks, we predefine a low-dimensional subspace to confine the activations and gradients, allowing full reconstruction in subsequent layers. Our method achieves up to 100x improvement in communication efficiency and enables training billion-parameter-scale models over low-end GPUs connected via consumer-grade internet speeds as low as 80Mbps, matching the convergence of centralized datacenter systems with 100Gbps connections with model parallel.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01260
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Ramasinghe, Sameera
Ajanthan, Thalaiyasingam
Avraham, Gil
Zuo, Yan
Long, Alexander
Machine Learning
Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks. While existing compression techniques are effective in data-parallel, they do not extend to model parallelism. Unlike data-parallel training, where weight gradients are exchanged, model-parallel requires compressing activations and activation gradients as they propagate through layers, accumulating compression errors. We propose a novel compression algorithm that compresses both forward and backward passes, enabling up to 99% compression with no convergence degradation with negligible memory/compute overhead. By leveraging a recursive structure in transformer networks, we predefine a low-dimensional subspace to confine the activations and gradients, allowing full reconstruction in subsequent layers. Our method achieves up to 100x improvement in communication efficiency and enables training billion-parameter-scale models over low-end GPUs connected via consumer-grade internet speeds as low as 80Mbps, matching the convergence of centralized datacenter systems with 100Gbps connections with model parallel.
title Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism
topic Machine Learning
url https://arxiv.org/abs/2506.01260