Salvato in:
Dettagli Bibliografici
Autori principali: Gu, Yida, Wang, Fakang, Fu, Jianhao, Sun, Zhenhang, Zhang, Qianyu, Zhao, Hairui, Liu, Xingchen, Tian, Yang, Huang, Wenjing, Liu, Zedong, Chen, Yifan, Yang, Jinwu, Zhou, Yueyuan, Zhao, Qian, Li, Haoxu, Wang, Tao, Yu, Feng, Wang, Zhan, Tan, Guangming, Tao, Dingwen
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:https://arxiv.org/abs/2605.04478
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913093934645248
author Gu, Yida
Wang, Fakang
Fu, Jianhao
Sun, Zhenhang
Zhang, Qianyu
Zhao, Hairui
Liu, Xingchen
Tian, Yang
Huang, Wenjing
Liu, Zedong
Chen, Yifan
Yang, Jinwu
Zhou, Yueyuan
Zhao, Qian
Li, Haoxu
Wang, Tao
Yu, Feng
Wang, Zhan
Tan, Guangming
Tao, Dingwen
author_facet Gu, Yida
Wang, Fakang
Fu, Jianhao
Sun, Zhenhang
Zhang, Qianyu
Zhao, Hairui
Liu, Xingchen
Tian, Yang
Huang, Wenjing
Liu, Zedong
Chen, Yifan
Yang, Jinwu
Zhou, Yueyuan
Zhao, Qian
Li, Haoxu
Wang, Tao
Yu, Feng
Wang, Zhan
Tan, Guangming
Tao, Dingwen
contents As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes-substantially outperforming existing solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2605_04478
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
Gu, Yida
Wang, Fakang
Fu, Jianhao
Sun, Zhenhang
Zhang, Qianyu
Zhao, Hairui
Liu, Xingchen
Tian, Yang
Huang, Wenjing
Liu, Zedong
Chen, Yifan
Yang, Jinwu
Zhou, Yueyuan
Zhao, Qian
Li, Haoxu
Wang, Tao
Yu, Feng
Wang, Zhan
Tan, Guangming
Tao, Dingwen
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. These anomalies typically manifest as slow/hang communication, the most frequent and time-consuming category to diagnose. However, traditional diagnostic methods remain inaccurate and inefficient, frequently requiring hours or even days for root cause analysis. To address this, we propose CCL-D, a high-precision diagnostic system designed to detect and locate slow/hang anomalies in large-scale distributed training. CCL-D integrates a rank-level real-time probe with an intelligent decision analyzer. The probe measures cross-layer anomaly metrics using a lightweight distributed tracing framework to monitor communication traffic. The analyzer performs automated anomaly detection and root-cause location, precisely identifying the faulty GPU rank. Deployed on a 4,000-GPU cluster over one year, CCL-D achieved near-complete coverage of known slow/hang anomalies and pinpointed affected ranks within 6 minutes-substantially outperforming existing solutions.
title CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
url https://arxiv.org/abs/2605.04478