Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Guanbin, Xu, ZhenGuo, Li, Yuzhe, Bai, Youhui, Gong, Ping, Ruan, Chaoyi, Li, Cheng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912923056603136
author Xu, Guanbin
Xu, ZhenGuo
Li, Yuzhe
Bai, Youhui
Gong, Ping
Ruan, Chaoyi
Li, Cheng
author_facet Xu, Guanbin
Xu, ZhenGuo
Li, Yuzhe
Bai, Youhui
Gong, Ping
Ruan, Chaoyi
Li, Cheng
contents Overlapping communication with computation is crucial for distributed large-model training, yet optimizing it - especially when computation becomes the bottleneck-remains challenging. We present Lagom, a system that co-tunes communication parameters to balance resource usage between computation and communication. By introducing a unified cost model and a priority-based search algorithm, Lagom reduces optimization complexity from exponential to linear. Evaluations on high- and low-bandwidth GPU clusters show that Lagom achieves 1.07-1.33x and 1.03-1.27x speedup over NCCL and AutoCCL across diverse models and parallelizations.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20656
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
Xu, Guanbin
Xu, ZhenGuo
Li, Yuzhe
Bai, Youhui
Gong, Ping
Ruan, Chaoyi
Li, Cheng
Distributed, Parallel, and Cluster Computing
Overlapping communication with computation is crucial for distributed large-model training, yet optimizing it - especially when computation becomes the bottleneck-remains challenging. We present Lagom, a system that co-tunes communication parameters to balance resource usage between computation and communication. By introducing a unified cost model and a priority-based search algorithm, Lagom reduces optimization complexity from exponential to linear. Evaluations on high- and low-bandwidth GPU clusters show that Lagom achieves 1.07-1.33x and 1.03-1.27x speedup over NCCL and AutoCCL across diverse models and parallelizations.
title Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2602.20656