Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Xuan-Phi, Pandit, Shrey, Xu, Austin, Xiong, Caiming, Joty, Shafiq |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
von: Kiselev, O. M.
Veröffentlicht: (2026)
von: Kiselev, O. M.
Veröffentlicht: (2026)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
von: Shi, Zhenmei, et al.
Veröffentlicht: (2024)
SFR-RAG: Towards Contextually Faithful LLMs
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models
von: Han, X. Y., et al.
Veröffentlicht: (2025)
von: Han, X. Y., et al.
Veröffentlicht: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
von: Hu, Xing, et al.
Veröffentlicht: (2025)
von: Hu, Xing, et al.
Veröffentlicht: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
Mixture of A Million Experts
von: He, Xu Owen
Veröffentlicht: (2024)
von: He, Xu Owen
Veröffentlicht: (2024)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
von: Tran, Huyen Ngoc, et al.
Veröffentlicht: (2026)
von: Tran, Huyen Ngoc, et al.
Veröffentlicht: (2026)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
von: Yan, Jiaming, et al.
Veröffentlicht: (2025)
Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Mixture of Raytraced Experts
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
von: Perin, Andrea, et al.
Veröffentlicht: (2025)
Wavelet Mixture of Experts for Time Series Forecasting
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
von: Zhou, Zheng, et al.
Veröffentlicht: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
von: Gao, Yuting, et al.
Veröffentlicht: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
von: He, Yifei, et al.
Veröffentlicht: (2025)
von: He, Yifei, et al.
Veröffentlicht: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Load Balancing Mixture of Experts with Similarity Preserving Routers
von: Omi, Nabil, et al.
Veröffentlicht: (2025)
von: Omi, Nabil, et al.
Veröffentlicht: (2025)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
von: Zhao, Zirui, et al.
Veröffentlicht: (2024)
von: Zhao, Zirui, et al.
Veröffentlicht: (2024)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
von: Dwivedi, Chaitanya, et al.
Veröffentlicht: (2026)
MoNTA: Accelerating Mixture-of-Experts Training with Network-Traffc-Aware Parallel Optimization
von: Guo, Jingming, et al.
Veröffentlicht: (2024)
von: Guo, Jingming, et al.
Veröffentlicht: (2024)
Mixture of Concept Bottleneck Experts
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
von: De Santis, Francesco, et al.
Veröffentlicht: (2026)
Mixture of Diverse Size Experts
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
von: Sun, Manxi, et al.
Veröffentlicht: (2024)
Sparsity and Superposition in Mixture of Experts
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2025)
Mixture of Experts in a Mixture of RL settings
von: Willi, Timon, et al.
Veröffentlicht: (2024)
von: Willi, Timon, et al.
Veröffentlicht: (2024)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
von: Chu, Kexin, et al.
Veröffentlicht: (2025)
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
von: Wang, Lean, et al.
Veröffentlicht: (2024)
von: Wang, Lean, et al.
Veröffentlicht: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
MC#: Mixture Compressor for Mixture-of-Experts Large Models
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025) -
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024) -
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025) -
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
von: Pandit, Shrey, et al.
Veröffentlicht: (2025) -
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)