Robust LLM Training Infrastructure at ByteDance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Borui, Liu, Gaohong, Song, Zuquan, Wang, Jun, Zhang, Yun, Sheng, Guangming, Wang, Shuguang, Wei, Houmin, Wang, Chenyuan, Lou, Weiqiang, Yang, Xi, Zhang, Mofan, Jiang, Kaihua, Ren, Cheng, Zhi, Xiaoyun, Yu, Menghan, Nan, Zhe, Zheng, Zhuolin, Zhong, Baoquan, Wang, Qinlong, Yu, Huan, Chi, Jinxin, Zhang, Wang, Li, Yuhan, Du, Zixian, Zhao, Sida, Zhang, Yongqiang, Tang, Jingzhe, Liu, Zherui, Wu, Chuan, Peng, Yanghua, Lin, Haibin, Xiao, Wencong, Liu, Xin, Xiang, Liang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917026065285120
author Wan, Borui
Liu, Gaohong
Song, Zuquan
Wang, Jun
Zhang, Yun
Sheng, Guangming
Wang, Shuguang
Wei, Houmin
Wang, Chenyuan
Lou, Weiqiang
Yang, Xi
Zhang, Mofan
Jiang, Kaihua
Ren, Cheng
Zhi, Xiaoyun
Yu, Menghan
Nan, Zhe
Zheng, Zhuolin
Zhong, Baoquan
Wang, Qinlong
Yu, Huan
Chi, Jinxin
Zhang, Wang
Li, Yuhan
Du, Zixian
Zhao, Sida
Zhang, Yongqiang
Tang, Jingzhe
Liu, Zherui
Wu, Chuan
Peng, Yanghua
Lin, Haibin
Xiao, Wencong
Liu, Xin
Xiang, Liang
author_facet Wan, Borui
Liu, Gaohong
Song, Zuquan
Wang, Jun
Zhang, Yun
Sheng, Guangming
Wang, Shuguang
Wei, Houmin
Wang, Chenyuan
Lou, Weiqiang
Yang, Xi
Zhang, Mofan
Jiang, Kaihua
Ren, Cheng
Zhi, Xiaoyun
Yu, Menghan
Nan, Zhe
Zheng, Zhuolin
Zhong, Baoquan
Wang, Qinlong
Yu, Huan
Chi, Jinxin
Zhang, Wang
Li, Yuhan
Du, Zixian
Zhao, Sida
Zhang, Yongqiang
Tang, Jingzhe
Liu, Zherui
Wu, Chuan
Peng, Yanghua
Lin, Haibin
Xiao, Wencong
Liu, Xin
Xiang, Liang
contents The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform and achieves 97% ETTR for a three-month training job on 9,600 GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16293
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust LLM Training Infrastructure at ByteDance
Wan, Borui
Liu, Gaohong
Song, Zuquan
Wang, Jun
Zhang, Yun
Sheng, Guangming
Wang, Shuguang
Wei, Houmin
Wang, Chenyuan
Lou, Weiqiang
Yang, Xi
Zhang, Mofan
Jiang, Kaihua
Ren, Cheng
Zhi, Xiaoyun
Yu, Menghan
Nan, Zhe
Zheng, Zhuolin
Zhong, Baoquan
Wang, Qinlong
Yu, Huan
Chi, Jinxin
Zhang, Wang
Li, Yuhan
Du, Zixian
Zhao, Sida
Zhang, Yongqiang
Tang, Jingzhe
Liu, Zherui
Wu, Chuan
Peng, Yanghua
Lin, Haibin
Xiao, Wencong
Liu, Xin
Xiang, Liang
Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform and achieves 97% ETTR for a three-month training job on 9,600 GPUs.
title Robust LLM Training Infrastructure at ByteDance
topic Machine Learning
Artificial Intelligence
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.16293