Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866912923056603136 |
|---|---|
| author | Xu, Guanbin Xu, ZhenGuo Li, Yuzhe Bai, Youhui Gong, Ping Ruan, Chaoyi Li, Cheng |
| author_facet | Xu, Guanbin Xu, ZhenGuo Li, Yuzhe Bai, Youhui Gong, Ping Ruan, Chaoyi Li, Cheng |
| contents | Overlapping communication with computation is crucial for distributed large-model training, yet optimizing it - especially when computation becomes the bottleneck-remains challenging. We present Lagom, a system that co-tunes communication parameters to balance resource usage between computation and communication. By introducing a unified cost model and a priority-based search algorithm, Lagom reduces optimization complexity from exponential to linear. Evaluations on high- and low-bandwidth GPU clusters show that Lagom achieves 1.07-1.33x and 1.03-1.27x speedup over NCCL and AutoCCL across diverse models and parallelizations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_20656 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training Xu, Guanbin Xu, ZhenGuo Li, Yuzhe Bai, Youhui Gong, Ping Ruan, Chaoyi Li, Cheng Distributed, Parallel, and Cluster Computing Overlapping communication with computation is crucial for distributed large-model training, yet optimizing it - especially when computation becomes the bottleneck-remains challenging. We present Lagom, a system that co-tunes communication parameters to balance resource usage between computation and communication. By introducing a unified cost model and a priority-based search algorithm, Lagom reduces optimization complexity from exponential to linear. Evaluations on high- and low-bandwidth GPU clusters show that Lagom achieves 1.07-1.33x and 1.03-1.27x speedup over NCCL and AutoCCL across diverse models and parallelizations. |
| title | Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2602.20656 |