Understanding Stragglers in Large Model Training Using What-if Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Jinkun, Jiang, Ziheng, Song, Zuquan, Zhao, Sida, Yu, Menghan, Wang, Zhanghan, Wang, Chenyuan, Shi, Zuocheng, Shi, Xiang, Jia, Wei, Liu, Zherui, Wang, Shuguang, Lin, Haibin, Liu, Xin, Panda, Aurojit, Li, Jinyang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916733274554368
author Lin, Jinkun
Jiang, Ziheng
Song, Zuquan
Zhao, Sida
Yu, Menghan
Wang, Zhanghan
Wang, Chenyuan
Shi, Zuocheng
Shi, Xiang
Jia, Wei
Liu, Zherui
Wang, Shuguang
Lin, Haibin
Liu, Xin
Panda, Aurojit
Li, Jinyang
author_facet Lin, Jinkun
Jiang, Ziheng
Song, Zuquan
Zhao, Sida
Yu, Menghan
Wang, Zhanghan
Wang, Chenyuan
Shi, Zuocheng
Shi, Xiang
Jia, Wei
Liu, Zherui
Wang, Shuguang
Lin, Haibin
Liu, Xin
Panda, Aurojit
Li, Jinyang
contents Large language model (LLM) training is one of the most demanding distributed computations today, often requiring thousands of GPUs with frequent synchronization across machines. Such a workload pattern makes it susceptible to stragglers, where the training can be stalled by few slow workers. At ByteDance we find stragglers are not trivially always caused by hardware failures, but can arise from multiple complex factors. This work aims to present a comprehensive study on the straggler issues in LLM training, using a five-month trace collected from our ByteDance LLM training cluster. The core methodology is what-if analysis that simulates the scenario without any stragglers and contrasts with the actual case. We use this method to study the following questions: (1) how often do stragglers affect training jobs, and what effect do they have on job performance; (2) do stragglers exhibit temporal or spatial patterns; and (3) what are the potential root causes for stragglers?
format Preprint
id arxiv_https___arxiv_org_abs_2505_05713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding Stragglers in Large Model Training Using What-if Analysis
Lin, Jinkun
Jiang, Ziheng
Song, Zuquan
Zhao, Sida
Yu, Menghan
Wang, Zhanghan
Wang, Chenyuan
Shi, Zuocheng
Shi, Xiang
Jia, Wei
Liu, Zherui
Wang, Shuguang
Lin, Haibin
Liu, Xin
Panda, Aurojit
Li, Jinyang
Distributed, Parallel, and Cluster Computing
Machine Learning
Large language model (LLM) training is one of the most demanding distributed computations today, often requiring thousands of GPUs with frequent synchronization across machines. Such a workload pattern makes it susceptible to stragglers, where the training can be stalled by few slow workers. At ByteDance we find stragglers are not trivially always caused by hardware failures, but can arise from multiple complex factors. This work aims to present a comprehensive study on the straggler issues in LLM training, using a five-month trace collected from our ByteDance LLM training cluster. The core methodology is what-if analysis that simulates the scenario without any stragglers and contrasts with the actual case. We use this method to study the following questions: (1) how often do stragglers affect training jobs, and what effect do they have on job performance; (2) do stragglers exhibit temporal or spatial patterns; and (3) what are the potential root causes for stragglers?
title Understanding Stragglers in Large Model Training Using What-if Analysis
topic Distributed, Parallel, and Cluster Computing
Machine Learning
url https://arxiv.org/abs/2505.05713