X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918126771240960 |
|---|---|
| author | Yuan, Yueming Gupta, Ahan Li, Jianping Dash, Sajal Wang, Feiyi Zhang, Minjia |
| author_facet | Yuan, Yueming Gupta, Ahan Li, Jianping Dash, Sajal Wang, Feiyi Zhang, Minjia |
| contents | Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k routing. However, their scalability is limited by substantial activation memory overhead and costly all-to-all communication. Furthermore, current MoE training systems - primarily optimized for NVIDIA GPUs - perform suboptimally on non-NVIDIA platforms, leaving significant computational potential untapped. In this work, we present X-MoE, a novel MoE training system designed to deliver scalable training performance for next-generation MoE architectures. X-MoE achieves this via several novel techniques, including efficient padding-free MoE training with cross-platform kernels, redundancy-bypassing dispatch, and hybrid parallelism with sequence-sharded MoE blocks. Our evaluation on the Frontier supercomputer, powered by AMD MI250X GPUs, shows that X-MoE scales DeepSeek-style MoEs up to 545 billion parameters across 1024 GPUs - 10x larger than the largest trainable model with existing methods under the same hardware budget, while maintaining high training throughput. The source code of X-MoE is available at https://github.com/Supercomputing-System-AI-Lab/X-MoE. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_13337 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms Yuan, Yueming Gupta, Ahan Li, Jianping Dash, Sajal Wang, Feiyi Zhang, Minjia Machine Learning Computation and Language Distributed, Parallel, and Cluster Computing Emerging expert-specialized Mixture-of-Experts (MoE) architectures, such as DeepSeek-MoE, deliver strong model quality through fine-grained expert segmentation and large top-k routing. However, their scalability is limited by substantial activation memory overhead and costly all-to-all communication. Furthermore, current MoE training systems - primarily optimized for NVIDIA GPUs - perform suboptimally on non-NVIDIA platforms, leaving significant computational potential untapped. In this work, we present X-MoE, a novel MoE training system designed to deliver scalable training performance for next-generation MoE architectures. X-MoE achieves this via several novel techniques, including efficient padding-free MoE training with cross-platform kernels, redundancy-bypassing dispatch, and hybrid parallelism with sequence-sharded MoE blocks. Our evaluation on the Frontier supercomputer, powered by AMD MI250X GPUs, shows that X-MoE scales DeepSeek-style MoEs up to 545 billion parameters across 1024 GPUs - 10x larger than the largest trainable model with existing methods under the same hardware budget, while maintaining high training throughput. The source code of X-MoE is available at https://github.com/Supercomputing-System-AI-Lab/X-MoE. |
| title | X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms |
| topic | Machine Learning Computation and Language Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2508.13337 |