MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Iacob, Alex, Jovanovic, Andrej, Safaryan, Mher, Kurmanji, Meghdad, Sani, Lorenzo, Horváth, Samuel, Shen, William F., Qiu, Xinchi, Lane, Nicholas D.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915535682273280
author Iacob, Alex
Jovanovic, Andrej
Safaryan, Mher
Kurmanji, Meghdad
Sani, Lorenzo
Horváth, Samuel
Shen, William F.
Qiu, Xinchi
Lane, Nicholas D.
author_facet Iacob, Alex
Jovanovic, Andrej
Safaryan, Mher
Kurmanji, Meghdad
Sani, Lorenzo
Horváth, Samuel
Shen, William F.
Qiu, Xinchi
Lane, Nicholas D.
contents Training large models with distributed data parallelism (DDP) requires frequent communication of gradients across workers, which can saturate bandwidth. Infrequent communication strategies (e.g., Local SGD) reduce this overhead but, when applied to adaptive optimizers, often suffer a performance gap relative to fully synchronous DDP. We trace this gap to a time-scale mismatch: the optimizer's fast-moving momentum, tuned for frequent updates, decays too quickly to smooth gradients over long intervals, leading to noise-dominated optimization. To address this, we propose MT-DAO, a family of optimizers that employs multiple slow- and fast-moving first momenta or the gradient to track update dynamics across different time scales, for which we provide the first convergence guarantees. Empirically, for language-model pre-training, this eliminates the performance gap with DDP, outperforming infrequent-communication baselines in perplexity and reducing iso-token wall-clock time by 6-27% on Ethernet interconnects. At the 720M scale, MT-DAO reaches a target perplexity in 24% fewer steps and 35% less time than the single-momentum DDP baseline. MT-DAO enables effective cross-datacenter training and training over wide geographic areas.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
Iacob, Alex
Jovanovic, Andrej
Safaryan, Mher
Kurmanji, Meghdad
Sani, Lorenzo
Horváth, Samuel
Shen, William F.
Qiu, Xinchi
Lane, Nicholas D.
Machine Learning
Artificial Intelligence
Training large models with distributed data parallelism (DDP) requires frequent communication of gradients across workers, which can saturate bandwidth. Infrequent communication strategies (e.g., Local SGD) reduce this overhead but, when applied to adaptive optimizers, often suffer a performance gap relative to fully synchronous DDP. We trace this gap to a time-scale mismatch: the optimizer's fast-moving momentum, tuned for frequent updates, decays too quickly to smooth gradients over long intervals, leading to noise-dominated optimization. To address this, we propose MT-DAO, a family of optimizers that employs multiple slow- and fast-moving first momenta or the gradient to track update dynamics across different time scales, for which we provide the first convergence guarantees. Empirically, for language-model pre-training, this eliminates the performance gap with DDP, outperforming infrequent-communication baselines in perplexity and reducing iso-token wall-clock time by 6-27% on Ethernet interconnects. At the 720M scale, MT-DAO reaches a target perplexity in 24% fewer steps and 35% less time than the single-momentum DDP baseline. MT-DAO enables effective cross-datacenter training and training over wide geographic areas.
title MT-DAO: Multi-Timescale Distributed Adaptive Optimizers with Local Updates
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.05361