Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Weilin, Han, Jingtao, Zhang, Weizhong, Jin, Cheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Computational Budget Should Be Considered in Data Selection
von: Wan, Weilin, et al.
Veröffentlicht: (2025)
von: Wan, Weilin, et al.
Veröffentlicht: (2025)
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
von: Wan, Weilin, et al.
Veröffentlicht: (2025)
von: Wan, Weilin, et al.
Veröffentlicht: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Scaling Laws for Optimal Data Mixtures
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
von: Abnar, Samira, et al.
Veröffentlicht: (2025)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
von: Mayaki, Mansour Zoubeirou a
Veröffentlicht: (2026)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
von: Li, Junzhuo, et al.
Veröffentlicht: (2026)
von: Li, Junzhuo, et al.
Veröffentlicht: (2026)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
von: Song, Chenyang, et al.
Veröffentlicht: (2026)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
On Linear Mode Connectivity of Mixture-of-Experts Architectures
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
von: Tran, Viet-Hoang, et al.
Veröffentlicht: (2025)
A Survey on Mixture of Experts in Large Language Models
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
MoFE: Mixture of Frozen Experts Architecture
von: Seo, Jean, et al.
Veröffentlicht: (2025)
von: Seo, Jean, et al.
Veröffentlicht: (2025)
What Scales in Cross-Entropy Scaling Law?
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
von: Yan, Junxi, et al.
Veröffentlicht: (2025)
Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
von: Jha, Nandan Kumar, et al.
Veröffentlicht: (2026)
Slicing and Dicing: Configuring Optimal Mixtures of Experts
von: Li, Margaret, et al.
Veröffentlicht: (2026)
von: Li, Margaret, et al.
Veröffentlicht: (2026)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
von: Song, Chenyang, et al.
Veröffentlicht: (2025)
Mixture of Online and Offline Experts for Non-stationary Time Series
von: Zhao, Zhilin, et al.
Veröffentlicht: (2022)
von: Zhao, Zhilin, et al.
Veröffentlicht: (2022)
Adaptive Semantic Communication for Wireless Image Transmission Leveraging Mixture-of-Experts Mechanism
von: Wan, Haowen, et al.
Veröffentlicht: (2026)
von: Wan, Haowen, et al.
Veröffentlicht: (2026)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
von: Huang, Minbin, et al.
Veröffentlicht: (2026)
Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization
von: Li, Jingwei, et al.
Veröffentlicht: (2026)
von: Li, Jingwei, et al.
Veröffentlicht: (2026)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
FaST: Efficient and Effective Long-Horizon Forecasting for Large-Scale Spatial-Temporal Graphs via Mixture-of-Experts
von: Zhao, Yiji, et al.
Veröffentlicht: (2026)
von: Zhao, Yiji, et al.
Veröffentlicht: (2026)
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
Dynamic Mixture-of-Experts for Incremental Graph Learning
von: Kong, Lecheng, et al.
Veröffentlicht: (2025)
von: Kong, Lecheng, et al.
Veröffentlicht: (2025)
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
von: Vankadara, Leena Chennuru, et al.
Veröffentlicht: (2026)
von: Vankadara, Leena Chennuru, et al.
Veröffentlicht: (2026)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
von: Obando-Ceron, Johan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Computational Budget Should Be Considered in Data Selection
von: Wan, Weilin, et al.
Veröffentlicht: (2025) -
Explore and Establish Synergistic Effects Between Weight Pruning and Coreset Selection in Neural Network Training
von: Wan, Weilin, et al.
Veröffentlicht: (2025) -
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025) -
Scaling Laws for Optimal Data Mixtures
von: Shukor, Mustafa, et al.
Veröffentlicht: (2025) -
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
von: Abnar, Samira, et al.
Veröffentlicht: (2025)