Gespeichert in:
| Hauptverfasser: | Liew, Seng Pei, Shinzato, Kenta, Dong, Yuyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.08215 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Laws for Upcycling Mixture-of-Experts Language Models
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
Reusing Overtrained Language Models Saturates Scaling
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
von: Zhang, Di, et al.
Veröffentlicht: (2025)
von: Zhang, Di, et al.
Veröffentlicht: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
Bayesian Mixture of Experts For Large Language Models
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
von: Pan, Bowen, et al.
Veröffentlicht: (2024)
MLP Fusion: Towards Efficient Fine-tuning of Dense and Mixture-of-Experts Language Models
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
von: Ai, Mengting, et al.
Veröffentlicht: (2023)
A Survey on Mixture of Experts in Large Language Models
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
HMoE: Heterogeneous Mixture of Experts for Language Modeling
von: Wang, An, et al.
Veröffentlicht: (2024)
von: Wang, An, et al.
Veröffentlicht: (2024)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
von: Yoon, Youngsik, et al.
Veröffentlicht: (2026)
von: Yoon, Youngsik, et al.
Veröffentlicht: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
A Closer Look into Mixture-of-Experts in Large Language Models
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Maximum Score Routing For Mixture-of-Experts
von: Dong, Bowen, et al.
Veröffentlicht: (2025)
von: Dong, Bowen, et al.
Veröffentlicht: (2025)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
von: Dong, Zican, et al.
Veröffentlicht: (2025)
von: Dong, Zican, et al.
Veröffentlicht: (2025)
MoSE: Mixture of Slimmable Experts for Efficient and Adaptive Language Models
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
von: Tastan, Nurbek, et al.
Veröffentlicht: (2026)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
von: Zhong, Zexuan, et al.
Veröffentlicht: (2024)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
von: Liu, Baihui, et al.
Veröffentlicht: (2026)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Geometry-Preserving Aggregation for Mixture-of-Experts Embedding Models
von: Kachuee, Sajjad, et al.
Veröffentlicht: (2026)
von: Kachuee, Sajjad, et al.
Veröffentlicht: (2026)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
von: Sun, Yuan
Veröffentlicht: (2025)
von: Sun, Yuan
Veröffentlicht: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
von: Takashiro, Shota, et al.
Veröffentlicht: (2026)
Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
von: He, Haoze, et al.
Veröffentlicht: (2026)
von: He, Haoze, et al.
Veröffentlicht: (2026)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Laws for Upcycling Mixture-of-Experts Language Models
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025) -
Reusing Overtrained Language Models Saturates Scaling
von: Liew, Seng Pei, et al.
Veröffentlicht: (2025) -
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
von: Zhang, Di, et al.
Veröffentlicht: (2025) -
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025) -
Bayesian Mixture of Experts For Large Language Models
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)