Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lim, Jihyun, Jo, Junhyuk, Ko, Chanhyeok, Go, Young Min, Hwa, Jimin, Lee, Sunwoo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917287204749312
author Lim, Jihyun
Jo, Junhyuk
Ko, Chanhyeok
Go, Young Min
Hwa, Jimin
Lee, Sunwoo
author_facet Lim, Jihyun
Jo, Junhyuk
Ko, Chanhyeok
Go, Young Min
Hwa, Jimin
Lee, Sunwoo
contents Most parallel neural network training methods assume homogeneous computing resources. For example, synchronous data-parallel SGD suffers from significant synchronization overhead under heterogeneous workloads, often forcing practitioners to rely only on the fastest devices (e.g., GPUs). In this work, we study local SGD for efficient parallel training on heterogeneous systems. We show that intentionally introducing bias in data sampling and model aggregation can effectively harmonize slower CPUs with faster GPUs. Our extensive empirical results demonstrate that a carefully controlled bias significantly accelerates local SGD while achieving comparable or even higher accuracy than synchronous SGD under the same epoch budget. For instance, our method trains ResNet20 on CIFAR-10 with 2 CPUs and 8 GPUs up to 32x faster than synchronous SGD, with nearly identical accuracy. These results provide practical insights into how to flexibly utilize diverse compute resources for deep learning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08540
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
Lim, Jihyun
Jo, Junhyuk
Ko, Chanhyeok
Go, Young Min
Hwa, Jimin
Lee, Sunwoo
Machine Learning
Most parallel neural network training methods assume homogeneous computing resources. For example, synchronous data-parallel SGD suffers from significant synchronization overhead under heterogeneous workloads, often forcing practitioners to rely only on the fastest devices (e.g., GPUs). In this work, we study local SGD for efficient parallel training on heterogeneous systems. We show that intentionally introducing bias in data sampling and model aggregation can effectively harmonize slower CPUs with faster GPUs. Our extensive empirical results demonstrate that a carefully controlled bias significantly accelerates local SGD while achieving comparable or even higher accuracy than synchronous SGD under the same epoch budget. For instance, our method trains ResNet20 on CIFAR-10 with 2 CPUs and 8 GPUs up to 32x faster than synchronous SGD, with nearly identical accuracy. These results provide practical insights into how to flexibly utilize diverse compute resources for deep learning.
title Biased Local SGD for Efficient Deep Learning on Heterogeneous Systems
topic Machine Learning
url https://arxiv.org/abs/2508.08540