Reliable and Resilient Collective Communication Library for LLM Training and Serving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Wei, Yu, Nengneng, Xiong, Sixian, Liu, Zaoxing
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915701790343168
author Wang, Wei
Yu, Nengneng
Xiong, Sixian
Liu, Zaoxing
author_facet Wang, Wei
Yu, Nengneng
Xiong, Sixian
Liu, Zaoxing
contents Modern ML training and inference now span tens to tens of thousands of GPUs, where network faults can waste 10--15\% of GPU hours due to slow recovery. Common network errors and link fluctuations trigger timeouts that often terminate entire jobs, forcing expensive checkpoint rollback during training and request reprocessing during inference. We present R$^2$CCL, a fault-tolerant communication library that provides lossless, low-overhead failover by exploiting multi-NIC hardware. R$^2$CCL performs rapid connection migration, bandwidth-aware load redistribution, and resilient collective algorithms to maintain progress under failures. We evaluate R$^2$CCL on two 8-GPU H100 InfiniBand servers and via large-scale ML simulators modeling hundreds of GPUs with diverse failure patterns. Experiments show that R$^2$CCL is highly robust to NIC failures, incurring less than 1\% training and less than 3\% inference overheads. R$^2$CCL outperforms baselines AdapCC and DejaVu by 12.18$\times$ and 47$\times$, respectively.
format Preprint
id arxiv_https___arxiv_org_abs_2512_25059
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reliable and Resilient Collective Communication Library for LLM Training and Serving
Wang, Wei
Yu, Nengneng
Xiong, Sixian
Liu, Zaoxing
Distributed, Parallel, and Cluster Computing
Machine Learning
Networking and Internet Architecture
Modern ML training and inference now span tens to tens of thousands of GPUs, where network faults can waste 10--15\% of GPU hours due to slow recovery. Common network errors and link fluctuations trigger timeouts that often terminate entire jobs, forcing expensive checkpoint rollback during training and request reprocessing during inference. We present R$^2$CCL, a fault-tolerant communication library that provides lossless, low-overhead failover by exploiting multi-NIC hardware. R$^2$CCL performs rapid connection migration, bandwidth-aware load redistribution, and resilient collective algorithms to maintain progress under failures. We evaluate R$^2$CCL on two 8-GPU H100 InfiniBand servers and via large-scale ML simulators modeling hundreds of GPUs with diverse failure patterns. Experiments show that R$^2$CCL is highly robust to NIC failures, incurring less than 1\% training and less than 3\% inference overheads. R$^2$CCL outperforms baselines AdapCC and DejaVu by 12.18$\times$ and 47$\times$, respectively.
title Reliable and Resilient Collective Communication Library for LLM Training and Serving
topic Distributed, Parallel, and Cluster Computing
Machine Learning
Networking and Internet Architecture
url https://arxiv.org/abs/2512.25059