Robust LLM Training Infrastructure at ByteDance
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917026065285120 |
|---|---|
| author | Wan, Borui Liu, Gaohong Song, Zuquan Wang, Jun Zhang, Yun Sheng, Guangming Wang, Shuguang Wei, Houmin Wang, Chenyuan Lou, Weiqiang Yang, Xi Zhang, Mofan Jiang, Kaihua Ren, Cheng Zhi, Xiaoyun Yu, Menghan Nan, Zhe Zheng, Zhuolin Zhong, Baoquan Wang, Qinlong Yu, Huan Chi, Jinxin Zhang, Wang Li, Yuhan Du, Zixian Zhao, Sida Zhang, Yongqiang Tang, Jingzhe Liu, Zherui Wu, Chuan Peng, Yanghua Lin, Haibin Xiao, Wencong Liu, Xin Xiang, Liang |
| author_facet | Wan, Borui Liu, Gaohong Song, Zuquan Wang, Jun Zhang, Yun Sheng, Guangming Wang, Shuguang Wei, Houmin Wang, Chenyuan Lou, Weiqiang Yang, Xi Zhang, Mofan Jiang, Kaihua Ren, Cheng Zhi, Xiaoyun Yu, Menghan Nan, Zhe Zheng, Zhuolin Zhong, Baoquan Wang, Qinlong Yu, Huan Chi, Jinxin Zhang, Wang Li, Yuhan Du, Zixian Zhao, Sida Zhang, Yongqiang Tang, Jingzhe Liu, Zherui Wu, Chuan Peng, Yanghua Lin, Haibin Xiao, Wencong Liu, Xin Xiang, Liang |
| contents | The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform and achieves 97% ETTR for a three-month training job on 9,600 GPUs. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_16293 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Robust LLM Training Infrastructure at ByteDance Wan, Borui Liu, Gaohong Song, Zuquan Wang, Jun Zhang, Yun Sheng, Guangming Wang, Shuguang Wei, Houmin Wang, Chenyuan Lou, Weiqiang Yang, Xi Zhang, Mofan Jiang, Kaihua Ren, Cheng Zhi, Xiaoyun Yu, Menghan Nan, Zhe Zheng, Zhuolin Zhong, Baoquan Wang, Qinlong Yu, Huan Chi, Jinxin Zhang, Wang Li, Yuhan Du, Zixian Zhao, Sida Zhang, Yongqiang Tang, Jingzhe Liu, Zherui Wu, Chuan Peng, Yanghua Lin, Haibin Xiao, Wencong Liu, Xin Xiang, Liang Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform and achieves 97% ETTR for a three-month training job on 9,600 GPUs. |
| title | Robust LLM Training Infrastructure at ByteDance |
| topic | Machine Learning Artificial Intelligence Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2509.16293 |