Understanding Stragglers in Large Model Training Using What-if Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916733274554368 |
|---|---|
| author | Lin, Jinkun Jiang, Ziheng Song, Zuquan Zhao, Sida Yu, Menghan Wang, Zhanghan Wang, Chenyuan Shi, Zuocheng Shi, Xiang Jia, Wei Liu, Zherui Wang, Shuguang Lin, Haibin Liu, Xin Panda, Aurojit Li, Jinyang |
| author_facet | Lin, Jinkun Jiang, Ziheng Song, Zuquan Zhao, Sida Yu, Menghan Wang, Zhanghan Wang, Chenyuan Shi, Zuocheng Shi, Xiang Jia, Wei Liu, Zherui Wang, Shuguang Lin, Haibin Liu, Xin Panda, Aurojit Li, Jinyang |
| contents | Large language model (LLM) training is one of the most demanding distributed computations today, often requiring thousands of GPUs with frequent synchronization across machines. Such a workload pattern makes it susceptible to stragglers, where the training can be stalled by few slow workers. At ByteDance we find stragglers are not trivially always caused by hardware failures, but can arise from multiple complex factors. This work aims to present a comprehensive study on the straggler issues in LLM training, using a five-month trace collected from our ByteDance LLM training cluster. The core methodology is what-if analysis that simulates the scenario without any stragglers and contrasts with the actual case. We use this method to study the following questions: (1) how often do stragglers affect training jobs, and what effect do they have on job performance; (2) do stragglers exhibit temporal or spatial patterns; and (3) what are the potential root causes for stragglers? |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_05713 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Understanding Stragglers in Large Model Training Using What-if Analysis Lin, Jinkun Jiang, Ziheng Song, Zuquan Zhao, Sida Yu, Menghan Wang, Zhanghan Wang, Chenyuan Shi, Zuocheng Shi, Xiang Jia, Wei Liu, Zherui Wang, Shuguang Lin, Haibin Liu, Xin Panda, Aurojit Li, Jinyang Distributed, Parallel, and Cluster Computing Machine Learning Large language model (LLM) training is one of the most demanding distributed computations today, often requiring thousands of GPUs with frequent synchronization across machines. Such a workload pattern makes it susceptible to stragglers, where the training can be stalled by few slow workers. At ByteDance we find stragglers are not trivially always caused by hardware failures, but can arise from multiple complex factors. This work aims to present a comprehensive study on the straggler issues in LLM training, using a five-month trace collected from our ByteDance LLM training cluster. The core methodology is what-if analysis that simulates the scenario without any stragglers and contrasts with the actual case. We use this method to study the following questions: (1) how often do stragglers affect training jobs, and what effect do they have on job performance; (2) do stragglers exhibit temporal or spatial patterns; and (3) what are the potential root causes for stragglers? |
| title | Understanding Stragglers in Large Model Training Using What-if Analysis |
| topic | Distributed, Parallel, and Cluster Computing Machine Learning |
| url | https://arxiv.org/abs/2505.05713 |