Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tyagi, Sahil, Sharma, Prateek
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:https://arxiv.org/abs/2503.17469
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916659443269632
author Tyagi, Sahil
Sharma, Prateek
author_facet Tyagi, Sahil
Sharma, Prateek
contents Deep learning systems are optimized for clusters with homogeneous resources. However, heterogeneity is prevalent in computing infrastructure across edge, cloud and HPC. When training neural networks using stochastic gradient descent techniques on heterogeneous resources, performance degrades due to stragglers and stale updates. In this work, we develop an adaptive batch-scaling framework called OmniLearn to mitigate the effects of heterogeneity in distributed training. Our approach is inspired by proportional controllers to balance computation across heterogeneous servers, and works under varying resource availability. By dynamically adjusting worker mini-batches at runtime, OmniLearn reduces training time by 14-85%. We also investigate asynchronous training, where our techniques improve accuracy by up to 6.9%.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17469
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OmniLearn: A Framework for Distributed Deep Learning over Heterogeneous Clusters
Tyagi, Sahil
Sharma, Prateek
Machine Learning
Deep learning systems are optimized for clusters with homogeneous resources. However, heterogeneity is prevalent in computing infrastructure across edge, cloud and HPC. When training neural networks using stochastic gradient descent techniques on heterogeneous resources, performance degrades due to stragglers and stale updates. In this work, we develop an adaptive batch-scaling framework called OmniLearn to mitigate the effects of heterogeneity in distributed training. Our approach is inspired by proportional controllers to balance computation across heterogeneous servers, and works under varying resource availability. By dynamically adjusting worker mini-batches at runtime, OmniLearn reduces training time by 14-85%. We also investigate asynchronous training, where our techniques improve accuracy by up to 6.9%.
title OmniLearn: A Framework for Distributed Deep Learning over Heterogeneous Clusters
topic Machine Learning
url https://arxiv.org/abs/2503.17469