Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Abnar, Samira, Shah, Harshay, Busbridge, Dan, Ali, Alaaeldin Mohamed Elnouby, Susskind, Josh, Thilak, Vimal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
How PARTs assemble into wholes: Learning the relative composition of images
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Path-Constrained Mixture-of-Experts
by: Gu, Zijin, et al.
Published: (2026)
by: Gu, Zijin, et al.
Published: (2026)
Scaling Laws for Native Multimodal Models
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
by: Nakamura, Taishi, et al.
Published: (2025)
by: Nakamura, Taishi, et al.
Published: (2025)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
by: Li, Xianhang, et al.
Published: (2025)
by: Li, Xianhang, et al.
Published: (2025)
Vanishing Gradients in Reinforcement Finetuning of Language Models
by: Razin, Noam, et al.
Published: (2023)
by: Razin, Noam, et al.
Published: (2023)
Decomposing and Editing Predictions by Modeling Model Computation
by: Shah, Harshay, et al.
Published: (2024)
by: Shah, Harshay, et al.
Published: (2024)
Distillation Scaling Laws
by: Busbridge, Dan, et al.
Published: (2025)
by: Busbridge, Dan, et al.
Published: (2025)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Retrieval Capabilities of Large Language Models Scale with Pretraining FLOPs
by: Portes, Jacob, et al.
Published: (2025)
by: Portes, Jacob, et al.
Published: (2025)
Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization
by: Wan, Weilin, et al.
Published: (2026)
by: Wan, Weilin, et al.
Published: (2026)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
by: Littwin, Etai, et al.
Published: (2024)
by: Littwin, Etai, et al.
Published: (2024)
Revisiting the Scaling Properties of Downstream Metrics in Large Language Model Training
by: Krajewski, Jakub, et al.
Published: (2025)
by: Krajewski, Jakub, et al.
Published: (2025)
Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
by: Bethune, Louis, et al.
Published: (2025)
by: Bethune, Louis, et al.
Published: (2025)
Sparsity and Superposition in Mixture of Experts
by: Chaudhari, Marmik, et al.
Published: (2025)
by: Chaudhari, Marmik, et al.
Published: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
by: Tian, Changxin, et al.
Published: (2025)
by: Tian, Changxin, et al.
Published: (2025)
ContextCite: Attributing Model Generation to Context
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
by: Cohen-Wang, Benjamin, et al.
Published: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
Generalization and Scaling Laws for Mixture-of-Experts Transformers
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
by: Mayaki, Mansour Zoubeirou a
Published: (2026)
Compression Scaling Laws:Unifying Sparsity and Quantization
by: Frantar, Elias, et al.
Published: (2025)
by: Frantar, Elias, et al.
Published: (2025)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
by: Li, Junzhuo, et al.
Published: (2026)
by: Li, Junzhuo, et al.
Published: (2026)
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
by: Zhao, Jinze, et al.
Published: (2024)
by: Zhao, Jinze, et al.
Published: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
by: Zhao, Guoliang, et al.
Published: (2025)
by: Zhao, Guoliang, et al.
Published: (2025)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
by: Wu, Haoyuan, et al.
Published: (2024)
by: Wu, Haoyuan, et al.
Published: (2024)
Mixtures of Experts Unlock Parameter Scaling for Deep RL
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
by: Liu, Zehua, et al.
Published: (2025)
by: Liu, Zehua, et al.
Published: (2025)
Scaling Properties of Continuous Diffusion Spoken Language Models
by: Ramapuram, Jason, et al.
Published: (2026)
by: Ramapuram, Jason, et al.
Published: (2026)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts Conversion
by: Szatkowski, Filip, et al.
Published: (2023)
by: Szatkowski, Filip, et al.
Published: (2023)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
by: Tang, Yehui, et al.
Published: (2025)
by: Tang, Yehui, et al.
Published: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
by: Ludziejewski, Jan, et al.
Published: (2025)
by: Ludziejewski, Jan, et al.
Published: (2025)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
by: Thiombiano, Abdoul Majid O., et al.
Published: (2025)
Flopping for FLOPs: Leveraging equivariance for computational efficiency
by: Bökman, Georg, et al.
Published: (2025)
by: Bökman, Georg, et al.
Published: (2025)
Similar Items
-
Scaling Laws for Optimal Data Mixtures
by: Shukor, Mustafa, et al.
Published: (2025) -
How PARTs assemble into wholes: Learning the relative composition of images
by: Ayoughi, Melika, et al.
Published: (2025) -
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024) -
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026) -
Path-Constrained Mixture-of-Experts
by: Gu, Zijin, et al.
Published: (2026)