Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dong, Jianbo, Luo, Bin, Zhang, Jun, Zhang, Pengcheng, Feng, Fei, Zhu, Yikai, Liu, Ang, Chen, Zian, Shi, Yi, Jiao, Hairong, Lu, Gang, Guan, Yu, Zhai, Ennan, Xiao, Wencong, Zhao, Hanyu, Yuan, Man, Yang, Siran, Li, Xiang, Wang, Jiamang, Men, Rui, Zhang, Jianwei, Zhou, Chang, Cai, Dennis, Xie, Yuan, Fu, Binzhang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2406.04594
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913853613277184
author Dong, Jianbo
Luo, Bin
Zhang, Jun
Zhang, Pengcheng
Feng, Fei
Zhu, Yikai
Liu, Ang
Chen, Zian
Shi, Yi
Jiao, Hairong
Lu, Gang
Guan, Yu
Zhai, Ennan
Xiao, Wencong
Zhao, Hanyu
Yuan, Man
Yang, Siran
Li, Xiang
Wang, Jiamang
Men, Rui
Zhang, Jianwei
Zhou, Chang
Cai, Dennis
Xie, Yuan
Fu, Binzhang
author_facet Dong, Jianbo
Luo, Bin
Zhang, Jun
Zhang, Pengcheng
Feng, Fei
Zhu, Yikai
Liu, Ang
Chen, Zian
Shi, Yi
Jiao, Hairong
Lu, Gang
Guan, Yu
Zhai, Ennan
Xiao, Wencong
Zhao, Hanyu
Yuan, Man
Yang, Siran
Li, Xiang
Wang, Jiamang
Men, Rui
Zhang, Jianwei
Zhou, Chang
Cai, Dennis
Xie, Yuan
Fu, Binzhang
contents The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Moreover, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. And, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C4. The key insights of C4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, C4 can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The C4 has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to 45%. This enhancement is attributed to a 30% reduction in error-induced overhead and a 15% reduction in communication costs.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04594
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
Dong, Jianbo
Luo, Bin
Zhang, Jun
Zhang, Pengcheng
Feng, Fei
Zhu, Yikai
Liu, Ang
Chen, Zian
Shi, Yi
Jiao, Hairong
Lu, Gang
Guan, Yu
Zhai, Ennan
Xiao, Wencong
Zhao, Hanyu
Yuan, Man
Yang, Siran
Li, Xiang
Wang, Jiamang
Men, Rui
Zhang, Jianwei
Zhou, Chang
Cai, Dennis
Xie, Yuan
Fu, Binzhang
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased likelihood of hardware errors in high-end GPU products and the heightened risk of network traffic collisions. Moreover, any local hardware failure can disrupt training tasks, and the inability to swiftly identify faulty components leads to a significant waste of GPU resources. And, prolonged communication due to traffic collisions can substantially increase GPU waiting times. To address these challenges, we propose a communication-driven solution, namely the C4. The key insights of C4 are twofold. First, the load in distributed training exhibits homogeneous characteristics and is divided into iterations through periodic synchronization, therefore hardware anomalies would incur certain syndrome in collective communication. By leveraging this feature, C4 can rapidly identify the faulty components, swiftly isolate the anomaly, and restart the task, thereby avoiding resource wastage caused by delays in anomaly detection. Second, the predictable communication model of collective communication, involving a limited number of long-lived flows, allows C4 to efficiently execute traffic planning, substantially reducing bandwidth competition among these flows. The C4 has been extensively deployed across real-world production systems in a hyperscale cloud provider, yielding a significant improvement in system efficiency, from 30% to 45%. This enhancement is attributed to a 30% reduction in error-induced overhead and a 15% reduction in communication costs.
title Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2406.04594