Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Shigang, Ben-Nun, Tal, Nadiradze, Giorgi, Di Girolamo, Salvatore, Dryden, Nikoli, Alistarh, Dan, Hoefler, Torsten
Formato: Preprint
Publicado: 2020
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918127999123456
author Li, Shigang
Ben-Nun, Tal
Nadiradze, Giorgi
Di Girolamo, Salvatore
Dryden, Nikoli
Alistarh, Dan
Hoefler, Torsten
author_facet Li, Shigang
Ben-Nun, Tal
Nadiradze, Giorgi
Di Girolamo, Salvatore
Dryden, Nikoli
Alistarh, Dan
Hoefler, Torsten
contents Deep learning at scale is dominated by communication time. Distributing samples across nodes usually yields the best performance, but poses scaling challenges due to global information dissemination and load imbalance across uneven sample lengths. State-of-the-art decentralized optimizers mitigate the problem, but require more iterations to achieve the same accuracy as their globally-communicating counterparts. We present Wait-Avoiding Group Model Averaging (WAGMA) SGD, a wait-avoiding stochastic optimizer that reduces global communication via subgroup weight exchange. The key insight is a combination of algorithmic changes to the averaging scheme and the use of a group allreduce operation. We prove the convergence of WAGMA-SGD, and empirically show that it retains convergence rates similar to Allreduce-SGD. For evaluation, we train ResNet-50 on ImageNet; Transformer for machine translation; and deep reinforcement learning for navigation at scale. Compared with state-of-the-art decentralized SGD variants, WAGMA-SGD significantly improves training throughput (e.g., 2.1x on 1,024 GPUs for reinforcement learning), and achieves the fastest time-to-solution (e.g., the highest score using the shortest training time for Transformer).
format Preprint
id arxiv_https___arxiv_org_abs_2005_00124
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
Li, Shigang
Ben-Nun, Tal
Nadiradze, Giorgi
Di Girolamo, Salvatore
Dryden, Nikoli
Alistarh, Dan
Hoefler, Torsten
Distributed, Parallel, and Cluster Computing
Machine Learning
C.1.4; D.1.3; I.2
Deep learning at scale is dominated by communication time. Distributing samples across nodes usually yields the best performance, but poses scaling challenges due to global information dissemination and load imbalance across uneven sample lengths. State-of-the-art decentralized optimizers mitigate the problem, but require more iterations to achieve the same accuracy as their globally-communicating counterparts. We present Wait-Avoiding Group Model Averaging (WAGMA) SGD, a wait-avoiding stochastic optimizer that reduces global communication via subgroup weight exchange. The key insight is a combination of algorithmic changes to the averaging scheme and the use of a group allreduce operation. We prove the convergence of WAGMA-SGD, and empirically show that it retains convergence rates similar to Allreduce-SGD. For evaluation, we train ResNet-50 on ImageNet; Transformer for machine translation; and deep reinforcement learning for navigation at scale. Compared with state-of-the-art decentralized SGD variants, WAGMA-SGD significantly improves training throughput (e.g., 2.1x on 1,024 GPUs for reinforcement learning), and achieves the fastest time-to-solution (e.g., the highest score using the shortest training time for Transformer).
title Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging
topic Distributed, Parallel, and Cluster Computing
Machine Learning
C.1.4; D.1.3; I.2
url https://arxiv.org/abs/2005.00124