Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.04478 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913093934645248 |
|---|---|
| author | Gu, Yida Wang, Fakang Fu, Jianhao Sun, Zhenhang Zhang, Qianyu Zhao, Hairui Liu, Xingchen Tian, Yang Huang, Wenjing Liu, Zedong Chen, Yifan Yang, Jinwu Zhou, Yueyuan Zhao, Qian Li, Haoxu Wang, Tao Yu, Feng Wang, Zhan Tan, Guangming Tao, Dingwen |
| author_facet | Gu, Yida Wang, Fakang Fu, Jianhao Sun, Zhenhang Zhang, Qianyu Zhao, Hairui Liu, Xingchen Tian, Yang Huang, Wenjing Liu, Zedong Chen, Yifan Yang, Jinwu Zhou, Yueyuan Zhao, Qian Li, Haoxu Wang, Tao Yu, Feng Wang, Zhan Tan, Guangming Tao, Dingwen |
| contents | As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes-substantially outperforming existing solutions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_04478 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training Gu, Yida Wang, Fakang Fu, Jianhao Sun, Zhenhang Zhang, Qianyu Zhao, Hairui Liu, Xingchen Tian, Yang Huang, Wenjing Liu, Zedong Chen, Yifan Yang, Jinwu Zhou, Yueyuan Zhao, Qian Li, Haoxu Wang, Tao Yu, Feng Wang, Zhan Tan, Guangming Tao, Dingwen Distributed, Parallel, and Cluster Computing Artificial Intelligence As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes-substantially outperforming existing solutions. |
| title | CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training |
| topic | Distributed, Parallel, and Cluster Computing Artificial Intelligence |
| url | https://arxiv.org/abs/2605.04478 |