ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nabli, Adel, Fournier, Louis, Erbacher, Pierre, Serrano, Louis, Belilovsky, Eugene, Oyallon, Edouard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914090896588800
author Nabli, Adel
Fournier, Louis
Erbacher, Pierre
Serrano, Louis
Belilovsky, Eugene
Oyallon, Edouard
author_facet Nabli, Adel
Fournier, Louis
Erbacher, Pierre
Serrano, Louis
Belilovsky, Eugene
Oyallon, Edouard
contents Training LLMs relies on distributed implementations using multiple GPUs to compute gradients in parallel with sharded optimizers. However, synchronizing gradients in data parallel setups introduces communication overhead that grows with the number of workers, limiting parallelization efficiency. Local optimization algorithms reduce communications but incur high memory costs as they prevent optimizer state sharding, hindering scalability. To address this, we propose \textbf{AC}cumulate while \textbf{CO}mmunicate (ACCO), a memory-efficient optimization algorithm for distributed LLM training. By synchronizing delayed gradients while computing new ones, ACCO reduces GPU idle time and supports heterogeneous hardware. To mitigate the convergence issues caused by delayed updates, we introduce a novel technique ensuring training dynamics align with standard distributed optimization. Compared to ZeRO-1, our approach is significantly faster and scales effectively across heterogeneous hardware.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
Nabli, Adel
Fournier, Louis
Erbacher, Pierre
Serrano, Louis
Belilovsky, Eugene
Oyallon, Edouard
Machine Learning
Artificial Intelligence
Training LLMs relies on distributed implementations using multiple GPUs to compute gradients in parallel with sharded optimizers. However, synchronizing gradients in data parallel setups introduces communication overhead that grows with the number of workers, limiting parallelization efficiency. Local optimization algorithms reduce communications but incur high memory costs as they prevent optimizer state sharding, hindering scalability. To address this, we propose \textbf{AC}cumulate while \textbf{CO}mmunicate (ACCO), a memory-efficient optimization algorithm for distributed LLM training. By synchronizing delayed gradients while computing new ones, ACCO reduces GPU idle time and supports heterogeneous hardware. To mitigate the convergence issues caused by delayed updates, we introduce a novel technique ensuring training dynamics align with standard distributed optimization. Compared to ZeRO-1, our approach is significantly faster and scales effectively across heterogeneous hardware.
title ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.02613