Grouter: Decoupling Routing from Representation for Accelerated MoE Training

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xu, Yuqi, Hu, Rizhen, Liu, Zihan, Sun, Mou, Yuan, Kun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917527555145728
author Xu, Yuqi
Hu, Rizhen
Liu, Zihan
Sun, Mou
Yuan, Kun
author_facet Xu, Yuqi
Hu, Rizhen
Liu, Zihan
Sun, Mou
Yuan, Kun
contents Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searching for an optimal routing policy within a vast combinatorial space. This entanglement often leads to sluggish convergence and training instabilities. This paper introduces Grouter, a preemptive routing method that by distilling high-quality structures from fully-trained MoE models and serving as a fixed router for target models. By decoupling structural optimization from weight updates, Grouter significantly accelerates both the speed and quality of model convergence. To ensure the framework's versatility, we also introduce expert folding to adapt Grouter across varying model configurations and expert tuning to rebalance workloads across different data distributions. Furthermore, by leveraging the structural priors provided by preemptive routing, we can implement targeted optimizations to further enhance training throughput. Experiments demonstrate that Grouter achieves superior performance and efficiency which boosts pre-training data utilization by 4.28x and achieves up to 33.5% throughput acceleration, establishing preemptive routing as a fundamental paradigm for scalable MoE training. We publicly release our code and pretrained Grouter checkpoints at https://github.com/JimmyAwoe/Grouter.
format Preprint
id arxiv_https___arxiv_org_abs_2603_06626
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Xu, Yuqi
Hu, Rizhen
Liu, Zihan
Sun, Mou
Yuan, Kun
Machine Learning
Artificial Intelligence
Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneously train expert weights while searching for an optimal routing policy within a vast combinatorial space. This entanglement often leads to sluggish convergence and training instabilities. This paper introduces Grouter, a preemptive routing method that by distilling high-quality structures from fully-trained MoE models and serving as a fixed router for target models. By decoupling structural optimization from weight updates, Grouter significantly accelerates both the speed and quality of model convergence. To ensure the framework's versatility, we also introduce expert folding to adapt Grouter across varying model configurations and expert tuning to rebalance workloads across different data distributions. Furthermore, by leveraging the structural priors provided by preemptive routing, we can implement targeted optimizations to further enhance training throughput. Experiments demonstrate that Grouter achieves superior performance and efficiency which boosts pre-training data utilization by 4.28x and achieves up to 33.5% throughput acceleration, establishing preemptive routing as a fundamental paradigm for scalable MoE training. We publicly release our code and pretrained Grouter checkpoints at https://github.com/JimmyAwoe/Grouter.
title Grouter: Decoupling Routing from Representation for Accelerated MoE Training
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.06626