Fine-grained MoE Load Balancing with Linear Programming

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Chenqi, Wu, Wenfei, Song, Linhai, Xu, Yuchen, Yuan, Yitao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912824581685248
author Zhao, Chenqi
Wu, Wenfei
Song, Linhai
Xu, Yuchen
Yuan, Yitao
author_facet Zhao, Chenqi
Wu, Wenfei
Song, Linhai
Xu, Yuchen
Yuan, Yitao
contents Mixture-of-Experts (MoE) has emerged as a promising approach to scale up deep learning models due to its significant reduction in computational resources. However, the dynamic nature of MoE leads to load imbalance among experts, severely impacting training efficiency. While previous research has attempted to address the load balancing challenge, existing solutions either compromise model accuracy or introduce additional system overhead. As a result, they fail to achieve fine-grained load balancing, which is crucial to optimizing training efficiency. We propose a novel parallelization strategy to achieve fine-grained load balancing in MoE systems. Our system is capable of achieving optimal load balancing in every micro-batch through efficient token scheduling across GPUs. Our experimental results demonstrate that MicroMoE improves the end-to-end training throughput by up to 47.6% compared with the state-of-the-art system, and almost consistently achieves optimal load balance among GPUs.
format Preprint
id arxiv_https___arxiv_org_abs_2511_16947
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-grained MoE Load Balancing with Linear Programming
Zhao, Chenqi
Wu, Wenfei
Song, Linhai
Xu, Yuchen
Yuan, Yitao
Distributed, Parallel, and Cluster Computing
Mixture-of-Experts (MoE) has emerged as a promising approach to scale up deep learning models due to its significant reduction in computational resources. However, the dynamic nature of MoE leads to load imbalance among experts, severely impacting training efficiency. While previous research has attempted to address the load balancing challenge, existing solutions either compromise model accuracy or introduce additional system overhead. As a result, they fail to achieve fine-grained load balancing, which is crucial to optimizing training efficiency. We propose a novel parallelization strategy to achieve fine-grained load balancing in MoE systems. Our system is capable of achieving optimal load balancing in every micro-batch through efficient token scheduling across GPUs. Our experimental results demonstrate that MicroMoE improves the end-to-end training throughput by up to 47.6% compared with the state-of-the-art system, and almost consistently achieves optimal load balance among GPUs.
title Fine-grained MoE Load Balancing with Linear Programming
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2511.16947