Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Guoliang, Jiang, Youhe, Xiao, Wencong, Jiang, Kaihua, Wang, Shuguang, Wang, Jun, Du, Zixian, Jiang, Zhuo, Zhang, Xinlei, Yuan, Binhang, Yoneki, Eiko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915502853455872
author He, Guoliang
Jiang, Youhe
Xiao, Wencong
Jiang, Kaihua
Wang, Shuguang
Wang, Jun
Du, Zixian
Jiang, Zhuo
Zhang, Xinlei
Yuan, Binhang
Yoneki, Eiko
author_facet He, Guoliang
Jiang, Youhe
Xiao, Wencong
Jiang, Kaihua
Wang, Shuguang
Wang, Jun
Du, Zixian
Jiang, Zhuo
Zhang, Xinlei
Yuan, Binhang
Yoneki, Eiko
contents The scaling law for large language models (LLMs) depicts that the path towards machine intelligence necessitates training at large scale. Thus, companies continuously build large-scale GPU clusters, and launch training jobs that span over thousands of computing nodes. However, LLM pre-training presents unique challenges due to its complex communication patterns, where GPUs exchange data in sparse yet high-volume bursts within specific groups. Inefficient resource scheduling exacerbates bandwidth contention, leading to suboptimal training performance. This paper presents Arnold, a scheduling system summarizing our experience to effectively align LLM communication patterns with data center topology at scale. An in-depth characteristic study is performed to identify the impact of physical network topology to LLM pre-training jobs. Based on the insights, we develop a scheduling algorithm to effectively align communication patterns with the physical network topology in modern data centers. Through simulation experiments, we show the effectiveness of our algorithm in reducing the maximum spread of communication groups by up to $1.67$x. In production training, our scheduling system improves the end-to-end performance by $10.6\%$ when training with more than $9600$ GPUs, a significant improvement for our training pipeline.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
He, Guoliang
Jiang, Youhe
Xiao, Wencong
Jiang, Kaihua
Wang, Shuguang
Wang, Jun
Du, Zixian
Jiang, Zhuo
Zhang, Xinlei
Yuan, Binhang
Yoneki, Eiko
Distributed, Parallel, and Cluster Computing
The scaling law for large language models (LLMs) depicts that the path towards machine intelligence necessitates training at large scale. Thus, companies continuously build large-scale GPU clusters, and launch training jobs that span over thousands of computing nodes. However, LLM pre-training presents unique challenges due to its complex communication patterns, where GPUs exchange data in sparse yet high-volume bursts within specific groups. Inefficient resource scheduling exacerbates bandwidth contention, leading to suboptimal training performance. This paper presents Arnold, a scheduling system summarizing our experience to effectively align LLM communication patterns with data center topology at scale. An in-depth characteristic study is performed to identify the impact of physical network topology to LLM pre-training jobs. Based on the insights, we develop a scheduling algorithm to effectively align communication patterns with the physical network topology in modern data centers. Through simulation experiments, we show the effectiveness of our algorithm in reducing the maximum spread of communication groups by up to $1.67$x. In production training, our scheduling system improves the end-to-end performance by $10.6\%$ when training with more than $9600$ GPUs, a significant improvement for our training pipeline.
title Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.15940