DiPaCo: Distributed Path Composition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Douillard, Arthur, Feng, Qixuan, Rusu, Andrei A., Kuncoro, Adhiguna, Donchev, Yani, Chhaparia, Rachita, Gog, Ionel, Ranzato, Marc'Aurelio, Shen, Jiajun, Szlam, Arthur
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909138267668480
author Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Kuncoro, Adhiguna
Donchev, Yani
Chhaparia, Rachita
Gog, Ionel
Ranzato, Marc'Aurelio
Shen, Jiajun
Szlam, Arthur
author_facet Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Kuncoro, Adhiguna
Donchev, Yani
Chhaparia, Rachita
Gog, Ionel
Ranzato, Marc'Aurelio
Shen, Jiajun
Szlam, Arthur
contents Progress in machine learning (ML) has been fueled by scaling neural network models. This scaling has been enabled by ever more heroic feats of engineering, necessary for accommodating ML approaches that require high bandwidth communication between devices working in parallel. In this work, we propose a co-designed modular architecture and training approach for ML models, dubbed DIstributed PAth COmposition (DiPaCo). During training, DiPaCo distributes computation by paths through a set of shared modules. Together with a Local-SGD inspired optimization (DiLoCo) that keeps modules in sync with drastically reduced communication, Our approach facilitates training across poorly connected and heterogeneous workers, with a design that ensures robustness to worker failures and preemptions. At inference time, only a single path needs to be executed for each input, without the need for any model compression. We consider this approach as a first prototype towards a new paradigm of large-scale learning, one that is less synchronous and more modular. Our experiments on the widely used C4 benchmark show that, for the same amount of training steps but less wall-clock time, DiPaCo exceeds the performance of a 1 billion-parameter dense transformer language model by choosing one of 256 possible paths, each with a size of 150 million parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10616
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiPaCo: Distributed Path Composition
Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Kuncoro, Adhiguna
Donchev, Yani
Chhaparia, Rachita
Gog, Ionel
Ranzato, Marc'Aurelio
Shen, Jiajun
Szlam, Arthur
Machine Learning
Computation and Language
Progress in machine learning (ML) has been fueled by scaling neural network models. This scaling has been enabled by ever more heroic feats of engineering, necessary for accommodating ML approaches that require high bandwidth communication between devices working in parallel. In this work, we propose a co-designed modular architecture and training approach for ML models, dubbed DIstributed PAth COmposition (DiPaCo). During training, DiPaCo distributes computation by paths through a set of shared modules. Together with a Local-SGD inspired optimization (DiLoCo) that keeps modules in sync with drastically reduced communication, Our approach facilitates training across poorly connected and heterogeneous workers, with a design that ensures robustness to worker failures and preemptions. At inference time, only a single path needs to be executed for each input, without the need for any model compression. We consider this approach as a first prototype towards a new paradigm of large-scale learning, one that is less synchronous and more modular. Our experiments on the widely used C4 benchmark show that, for the same amount of training steps but less wall-clock time, DiPaCo exceeds the performance of a 1 billion-parameter dense transformer language model by choosing one of 256 possible paths, each with a size of 150 million parameters.
title DiPaCo: Distributed Path Composition
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2403.10616