RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914108803121152 |
|---|---|
| author | Xu, Heng Yu, Zhiwei Du, Chengze Zhou, Ying Li, Letian Wang, Haojie Cheng, Weiqiang Li, Jialong |
| author_facet | Xu, Heng Yu, Zhiwei Du, Chengze Zhou, Ying Li, Letian Wang, Haojie Cheng, Weiqiang Li, Jialong |
| contents | Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology's symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%--78% and reduces completion time by 17%--78%. For Mixtral workloads, it shortens iteration time by 18%--40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_19262 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training Xu, Heng Yu, Zhiwei Du, Chengze Zhou, Ying Li, Letian Wang, Haojie Cheng, Weiqiang Li, Jialong Distributed, Parallel, and Cluster Computing Networking and Internet Architecture Training Mixture-of-Experts (MoE) models introduces sparse and highly imbalanced all-to-all communication that dominates iteration time. Conventional load-balancing methods fail to exploit the deterministic topology of Rail architectures, leaving multi-NIC bandwidth underutilized. We present RailS, a distributed load-balancing framework that minimizes all-to-all completion time in MoE training. RailS leverages the Rail topology's symmetry to prove that uniform sending ensures uniform receiving, transforming global coordination into local scheduling. Each node independently executes a Longest Processing Time First (LPT) spraying scheduler to proactively balance traffic using local information. RailS activates N parallel rails for fine-grained, topology-aware multipath transmission. Across synthetic and real-world MoE workloads, RailS improves bus bandwidth by 20%--78% and reduces completion time by 17%--78%. For Mixtral workloads, it shortens iteration time by 18%--40% and achieves near-optimal load balance, fully exploiting architectural parallelism in distributed training. |
| title | RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training |
| topic | Distributed, Parallel, and Cluster Computing Networking and Internet Architecture |
| url | https://arxiv.org/abs/2510.19262 |