Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Naganuma, Hiroki, Zhang, Xinzhi, Yue, Man-Chung, Mitliagkas, Ioannis, Witte, Philipp A., Hewett, Russell J., Lee, Yin Tat
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914040767315968
author Naganuma, Hiroki
Zhang, Xinzhi
Yue, Man-Chung
Mitliagkas, Ioannis
Witte, Philipp A.
Hewett, Russell J.
Lee, Yin Tat
author_facet Naganuma, Hiroki
Zhang, Xinzhi
Yue, Man-Chung
Mitliagkas, Ioannis
Witte, Philipp A.
Hewett, Russell J.
Lee, Yin Tat
contents Following AI scaling trends, frontier models continue to grow in size and continue to be trained on larger datasets. Training these models requires huge investments in exascale computational resources, which has in turn driven developtment of distributed deep learning methods. Data parallelism is an essential approach to speed up training, but it requires frequent global communication between workers, which can bottleneck training at the largest scales. In this work, we propose a method called Pseudo-Asynchronous Local SGD (PALSGD) to improve the efficiency of data-parallel training. PALSGD is an extension of Local SGD (Stich, 2018) and DiLoCo (Douillard et al., 2023), designed to further reduce communication frequency by introducing a pseudo-synchronization mechanism. PALSGD allows the use of longer synchronization intervals compared to standard Local SGD. Despite the reduced communication frequency, the pseudo-synchronization approach ensures that model consistency is maintained, leading to performance results comparable to those achieved with more frequent synchronization. Furthermore, we provide a theoretical analysis of PALSGD, establishing its convergence and deriving its convergence rate. This analysis offers insights into the algorithm's behavior and performance guarantees. We evaluated PALSGD on image classification and language modeling tasks. Our results show that PALSGD achieves better performance in less time compared to existing methods like Distributed Data Parallel (DDP), and DiLoCo. Notably, PALSGD trains 18.4% faster than DDP on ImageNet-1K with ResNet-50, 24.4% faster than DDP on TinyStories with GPT-Neo-125M, and 21.1% faster than DDP on TinyStories with GPT-Neo-8M.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
Naganuma, Hiroki
Zhang, Xinzhi
Yue, Man-Chung
Mitliagkas, Ioannis
Witte, Philipp A.
Hewett, Russell J.
Lee, Yin Tat
Machine Learning
Following AI scaling trends, frontier models continue to grow in size and continue to be trained on larger datasets. Training these models requires huge investments in exascale computational resources, which has in turn driven developtment of distributed deep learning methods. Data parallelism is an essential approach to speed up training, but it requires frequent global communication between workers, which can bottleneck training at the largest scales. In this work, we propose a method called Pseudo-Asynchronous Local SGD (PALSGD) to improve the efficiency of data-parallel training. PALSGD is an extension of Local SGD (Stich, 2018) and DiLoCo (Douillard et al., 2023), designed to further reduce communication frequency by introducing a pseudo-synchronization mechanism. PALSGD allows the use of longer synchronization intervals compared to standard Local SGD. Despite the reduced communication frequency, the pseudo-synchronization approach ensures that model consistency is maintained, leading to performance results comparable to those achieved with more frequent synchronization. Furthermore, we provide a theoretical analysis of PALSGD, establishing its convergence and deriving its convergence rate. This analysis offers insights into the algorithm's behavior and performance guarantees. We evaluated PALSGD on image classification and language modeling tasks. Our results show that PALSGD achieves better performance in less time compared to existing methods like Distributed Data Parallel (DDP), and DiLoCo. Notably, PALSGD trains 18.4% faster than DDP on ImageNet-1K with ResNet-50, 24.4% faster than DDP on TinyStories with GPT-Neo-125M, and 21.1% faster than DDP on TinyStories with GPT-Neo-8M.
title Pseudo-Asynchronous Local SGD: Robust and Efficient Data-Parallel Training
topic Machine Learning
url https://arxiv.org/abs/2504.18454