RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Heng, Yu, Zhiwei, Du, Chengze, Zhou, Ying, Li, Letian, Wang, Haojie, Cheng, Weiqiang, Li, Jialong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914108803121152
author Xu, Heng
Yu, Zhiwei
Du, Chengze
Zhou, Ying
Li, Letian
Wang, Haojie
Cheng, Weiqiang
Li, Jialong
author_facet Xu, Heng
Yu, Zhiwei
Du, Chengze
Zhou, Ying
Li, Letian
Wang, Haojie
Cheng, Weiqiang
Li, Jialong
contents Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology's symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%--78% and reduces completion time by 17%--78%. For Mixtral workloads, it shortens iteration time by 18%--40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
Xu, Heng
Yu, Zhiwei
Du, Chengze
Zhou, Ying
Li, Letian
Wang, Haojie
Cheng, Weiqiang
Li, Jialong
Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology's symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%--78% and reduces completion time by 17%--78%. For Mixtral workloads, it shortens iteration time by 18%--40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training.
title RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
topic Distributed, Parallel, and Cluster Computing
Networking and Internet Architecture
url https://arxiv.org/abs/2510.19262