Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Zihan, Huang, Zeyu, Zheng, Bo, Wen, Kaiyue, Wang, Zekun, Men, Rui, Titov, Ivan, Liu, Dayiheng, Zhou, Jingren, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
by: Qiu, Zihan, et al.
Published: (2026)
by: Qiu, Zihan, et al.
Published: (2026)
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024)
by: Qiu, Zihan, et al.
Published: (2024)
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)
by: Wang, Liangyu, et al.
Published: (2026)
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
by: Wang, Lean, et al.
Published: (2024)
by: Wang, Lean, et al.
Published: (2024)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
by: Tang, Shengkun, et al.
Published: (2026)
by: Tang, Shengkun, et al.
Published: (2026)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
by: Nguyen, Xuan-Phi, et al.
Published: (2026)
by: Nguyen, Xuan-Phi, et al.
Published: (2026)
Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training
by: Mouzouni, Charafeddine
Published: (2026)
by: Mouzouni, Charafeddine
Published: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
by: Xu, Heng, et al.
Published: (2025)
by: Xu, Heng, et al.
Published: (2025)
A Closer Look into Mixture-of-Experts in Large Language Models
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Load Balancing Mixture of Experts with Similarity Preserving Routers
by: Omi, Nabil, et al.
Published: (2025)
by: Omi, Nabil, et al.
Published: (2025)
$ϕ$-Balancing for Mixture-of-Experts Training
by: Chen, Lizhang, et al.
Published: (2026)
by: Chen, Lizhang, et al.
Published: (2026)
Post-hoc Reward Calibration: A Case Study on Length Bias
by: Huang, Zeyu, et al.
Published: (2024)
by: Huang, Zeyu, et al.
Published: (2024)
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
by: Jin, Haopeng
Published: (2026)
by: Jin, Haopeng
Published: (2026)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
by: Nikolic, Strahinja, et al.
Published: (2025)
by: Nikolic, Strahinja, et al.
Published: (2025)
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
by: Sun, Yuan
Published: (2025)
by: Sun, Yuan
Published: (2025)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
by: Han, X. Y., et al.
Published: (2025)
by: Han, X. Y., et al.
Published: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
by: Gritsch, Nikolas, et al.
Published: (2024)
by: Gritsch, Nikolas, et al.
Published: (2024)
Latent Prototype Routing: Achieving Near-Perfect Load Balancing in Mixture-of-Experts
by: Yang, Jiajie
Published: (2025)
by: Yang, Jiajie
Published: (2025)
Understanding and Leveraging the Expert Specialization of Context Faithfulness in Mixture-of-Experts LLMs
by: Bai, Jun, et al.
Published: (2025)
by: Bai, Jun, et al.
Published: (2025)
Fast Training of Mixture-of-Experts for Time Series Forecasting via Expert Loss Integration
by: Mahtout, Btissame El, et al.
Published: (2026)
by: Mahtout, Btissame El, et al.
Published: (2026)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Unlearning Traces the Influential Training Data of Language Models
by: Isonuma, Masaru, et al.
Published: (2024)
by: Isonuma, Masaru, et al.
Published: (2024)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
by: Liu, Xinyi, et al.
Published: (2026)
by: Liu, Xinyi, et al.
Published: (2026)
Temporally Extended Mixture-of-Experts Models
by: Shen, Zeyu, et al.
Published: (2026)
by: Shen, Zeyu, et al.
Published: (2026)
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
ProcessBench: Identifying Process Errors in Mathematical Reasoning
by: Zheng, Chujie, et al.
Published: (2024)
by: Zheng, Chujie, et al.
Published: (2024)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
by: Zhang, Zhenru, et al.
Published: (2025)
by: Zhang, Zhenru, et al.
Published: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
by: Huang, Zeyu, et al.
Published: (2025)
by: Huang, Zeyu, et al.
Published: (2025)
A Controllable Examination for Long-Context Language Models
by: Yang, Yijun, et al.
Published: (2025)
by: Yang, Yijun, et al.
Published: (2025)
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
by: Liu, Zijie, et al.
Published: (2026)
by: Liu, Zijie, et al.
Published: (2026)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
by: Zhao, Hao, et al.
Published: (2024)
by: Zhao, Hao, et al.
Published: (2024)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
by: Oldfield, James, et al.
Published: (2024)
by: Oldfield, James, et al.
Published: (2024)
The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits
by: Cheng, Tianhao, et al.
Published: (2026)
by: Cheng, Tianhao, et al.
Published: (2026)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
by: Lv, Ang, et al.
Published: (2025)
by: Lv, Ang, et al.
Published: (2025)
TriForecaster: A Mixture of Experts Framework for Multi-Region Electric Load Forecasting with Tri-dimensional Specialization
by: Zhu, Zhaoyang, et al.
Published: (2025)
by: Zhu, Zhaoyang, et al.
Published: (2025)
Extending the Limit Theorem of Barmpalias and Lewis-Pye to all reals
by: Titov, Ivan
Published: (2024)
by: Titov, Ivan
Published: (2024)
Variants of Solovay reducibility
by: Titov, Ivan
Published: (2024)
by: Titov, Ivan
Published: (2024)
Solovay reducibility implies S2a-reducibility
by: Titov, Ivan
Published: (2024)
by: Titov, Ivan
Published: (2024)
Similar Items
-
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025) -
A Unified View of Attention and Residual Sinks: Outlier-Driven Rescaling is Essential for Transformer Training
by: Qiu, Zihan, et al.
Published: (2026) -
Layerwise Recurrent Router for Mixture-of-Experts
by: Qiu, Zihan, et al.
Published: (2024) -
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026) -
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
by: Wang, Lean, et al.
Published: (2024)