DiLoCo: Distributed Low-Communication Training of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Douillard, Arthur, Feng, Qixuan, Rusu, Andrei A., Chhaparia, Rachita, Donchev, Yani, Kuncoro, Adhiguna, Ranzato, Marc'Aurelio, Szlam, Arthur, Shen, Jiajun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916406441803776
author Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Chhaparia, Rachita
Donchev, Yani
Kuncoro, Adhiguna
Ranzato, Marc'Aurelio
Szlam, Arthur
Shen, Jiajun
author_facet Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Chhaparia, Rachita
Donchev, Yani
Kuncoro, Adhiguna
Ranzato, Marc'Aurelio
Szlam, Arthur
Shen, Jiajun
contents Large language models (LLM) have become a critical component in many applications of machine learning. However, standard approaches to training LLM require a large number of tightly interconnected accelerators, with devices exchanging gradients and other intermediate states at each optimization step. While it is difficult to build and maintain a single computing cluster hosting many accelerators, it might be easier to find several computing clusters each hosting a smaller number of devices. In this work, we propose a distributed optimization algorithm, Distributed Low-Communication (DiLoCo), that enables training of language models on islands of devices that are poorly connected. The approach is a variant of federated averaging, where the number of inner steps is large, the inner optimizer is AdamW, and the outer optimizer is Nesterov momentum. On the widely used C4 dataset, we show that DiLoCo on 8 workers performs as well as fully synchronous optimization while communicating 500 times less. DiLoCo exhibits great robustness to the data distribution of each worker. It is also robust to resources becoming unavailable over time, and vice versa, it can seamlessly leverage resources that become available during training.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08105
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DiLoCo: Distributed Low-Communication Training of Language Models
Douillard, Arthur
Feng, Qixuan
Rusu, Andrei A.
Chhaparia, Rachita
Donchev, Yani
Kuncoro, Adhiguna
Ranzato, Marc'Aurelio
Szlam, Arthur
Shen, Jiajun
Machine Learning
Computation and Language
Large language models (LLM) have become a critical component in many applications of machine learning. However, standard approaches to training LLM require a large number of tightly interconnected accelerators, with devices exchanging gradients and other intermediate states at each optimization step. While it is difficult to build and maintain a single computing cluster hosting many accelerators, it might be easier to find several computing clusters each hosting a smaller number of devices. In this work, we propose a distributed optimization algorithm, Distributed Low-Communication (DiLoCo), that enables training of language models on islands of devices that are poorly connected. The approach is a variant of federated averaging, where the number of inner steps is large, the inner optimizer is AdamW, and the outer optimizer is Nesterov momentum. On the widely used C4 dataset, we show that DiLoCo on 8 workers performs as well as fully synchronous optimization while communicating 500 times less. DiLoCo exhibits great robustness to the data distribution of each worker. It is also robust to resources becoming unavailable over time, and vice versa, it can seamlessly leverage resources that become available during training.
title DiLoCo: Distributed Low-Communication Training of Language Models
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2311.08105