Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Weilin, Han, Jingtao, Zhang, Weizhong, Jin, Cheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Computational Budget Should Be Considered in Data Selection
by: Wan, Weilin, et al.
Published: (2025)
by: Wan, Weilin, et al.
Published: (2025)
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
by: Wan, Weilin, et al.
Published: (2025)
by: Wan, Weilin, et al.
Published: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025)
by: Zhao, Guoliang, et al.
Published: (2025)
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)
by: Abnar, Samira, et al.
Published: (2025)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
by: Li, Junzhuo, et al.
Published: (2026)
by: Li, Junzhuo, et al.
Published: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
by: Song, Chenyang, et al.
Published: (2026)
by: Song, Chenyang, et al.
Published: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
by: Zhu, Ruidong, et al.
Published: (2025)
by: Zhu, Ruidong, et al.
Published: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
by: He, Shwai, et al.
Published: (2024)
by: He, Shwai, et al.
Published: (2024)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
by: Tran, Viet-Hoang, et al.
Published: (2025)
by: Tran, Viet-Hoang, et al.
Published: (2025)
A Survey on Mixture of Experts in Large Language Models
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
by: He, Shwai, et al.
Published: (2025)
by: He, Shwai, et al.
Published: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
by: Liu, Yuzhi, et al.
Published: (2026)
by: Liu, Yuzhi, et al.
Published: (2026)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
by: Cai, Weilin, et al.
Published: (2024)
by: Cai, Weilin, et al.
Published: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
MoFE: Mixture of Frozen Experts Architecture
by: Seo, Jean, et al.
Published: (2025)
by: Seo, Jean, et al.
Published: (2025)
What Scales in Cross-Entropy Scaling Law?
by: Yan, Junxi, et al.
Published: (2025)
by: Yan, Junxi, et al.
Published: (2025)
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
by: Jha, Nandan Kumar, et al.
Published: (2026)
by: Jha, Nandan Kumar, et al.
Published: (2026)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
by: Li, Margaret, et al.
Published: (2026)
by: Li, Margaret, et al.
Published: (2026)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
by: Zheng, Chen, et al.
Published: (2025)
by: Zheng, Chen, et al.
Published: (2025)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
Mixture of Online and Offline Experts for Non-stationary Time Series
by: Zhao, Zhilin, et al.
Published: (2022)
by: Zhao, Zhilin, et al.
Published: (2022)
Adaptive Semantic Communication for Wireless Image Transmission Leveraging Mixture-of-Experts Mechanism
by: Wan, Haowen, et al.
Published: (2026)
by: Wan, Haowen, et al.
Published: (2026)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
by: Park, Sumin, et al.
Published: (2025)
by: Park, Sumin, et al.
Published: (2025)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
by: Aghdam, Maryam Akhavan, et al.
Published: (2024)
by: Aghdam, Maryam Akhavan, et al.
Published: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
by: Huang, Minbin, et al.
Published: (2026)
by: Huang, Minbin, et al.
Published: (2026)
Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization
by: Li, Jingwei, et al.
Published: (2026)
by: Li, Jingwei, et al.
Published: (2026)
Mixture of Lookup Experts
by: Jie, Shibo, et al.
Published: (2025)
by: Jie, Shibo, et al.
Published: (2025)
FaST: Efficient and Effective Long-Horizon Forecasting for Large-Scale Spatial-Temporal Graphs via Mixture-of-Experts
by: Zhao, Yiji, et al.
Published: (2026)
by: Zhao, Yiji, et al.
Published: (2026)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
by: Lv, Ang, et al.
Published: (2025)
by: Lv, Ang, et al.
Published: (2025)
Dynamic Mixture-of-Experts for Incremental Graph Learning
by: Kong, Lecheng, et al.
Published: (2025)
by: Kong, Lecheng, et al.
Published: (2025)
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
by: Vankadara, Leena Chennuru, et al.
Published: (2026)
by: Vankadara, Leena Chennuru, et al.
Published: (2026)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Similar Items
-
Computational Budget Should Be Considered in Data Selection
by: Wan, Weilin, et al.
Published: (2025) -
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
by: Wan, Weilin, et al.
Published: (2025) -
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025) -
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025) -
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
by: Abnar, Samira, et al.
Published: (2025)