Adaptive Consensus Gradients Aggregation for Scaled Distributed Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Choukroun, Yoni, Azoulay, Shlomi, Kisilev, Pavel
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910686684119040
author Choukroun, Yoni
Azoulay, Shlomi
Kisilev, Pavel
author_facet Choukroun, Yoni
Azoulay, Shlomi
Kisilev, Pavel
contents Distributed machine learning has recently become a critical paradigm for training large models on vast datasets. We examine the stochastic optimization problem for deep learning within synchronous parallel computing environments under communication constraints. While averaging distributed gradients is the most widely used method for gradient estimation, whether this is the optimal strategy remains an open question. In this work, we analyze the distributed gradient aggregation process through the lens of subspace optimization. By formulating the aggregation problem as an objective-aware subspace optimization problem, we derive an efficient weighting scheme for gradients, guided by subspace coefficients. We further introduce subspace momentum to accelerate convergence while maintaining statistical unbiasedness in the aggregation. Our method demonstrates improved performance over the ubiquitous gradient averaging on multiple MLPerf tasks while remaining extremely efficient in both communicational and computational complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03742
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
Choukroun, Yoni
Azoulay, Shlomi
Kisilev, Pavel
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
Distributed machine learning has recently become a critical paradigm for training large models on vast datasets. We examine the stochastic optimization problem for deep learning within synchronous parallel computing environments under communication constraints. While averaging distributed gradients is the most widely used method for gradient estimation, whether this is the optimal strategy remains an open question. In this work, we analyze the distributed gradient aggregation process through the lens of subspace optimization. By formulating the aggregation problem as an objective-aware subspace optimization problem, we derive an efficient weighting scheme for gradients, guided by subspace coefficients. We further introduce subspace momentum to accelerate convergence while maintaining statistical unbiasedness in the aggregation. Our method demonstrates improved performance over the ubiquitous gradient averaging on multiple MLPerf tasks while remaining extremely efficient in both communicational and computational complexity.
title Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2411.03742