Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Xudong, Liu, Qi, Xu, Yuhui, Zhou, Aojun, Huang, Siyuan, Zhang, Bo, Yan, Junchi, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
von: Guo, Hongcheng, et al.
Veröffentlicht: (2025)
von: Guo, Hongcheng, et al.
Veröffentlicht: (2025)
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
von: Huang, Yushi, et al.
Veröffentlicht: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
Unveiling Super Experts in Mixture-of-Experts Large Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
von: Su, Zunhai, et al.
Veröffentlicht: (2025)
Mixture Compressor for Mixture-of-Experts LLMs Gains More
von: Huang, Wei, et al.
Veröffentlicht: (2024)
von: Huang, Wei, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
von: Zhou, Yixiao, et al.
Veröffentlicht: (2025)
GatePro: Parameter-Free Expert Selection Optimization for Mixture-of-Experts Models
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
von: Zheng, Chen, et al.
Veröffentlicht: (2025)
DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
von: Bai, Sikai, et al.
Veröffentlicht: (2025)
von: Bai, Sikai, et al.
Veröffentlicht: (2025)
Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
von: Lv, Ang, et al.
Veröffentlicht: (2025)
von: Lv, Ang, et al.
Veröffentlicht: (2025)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
von: Jing, Linglin, et al.
Veröffentlicht: (2025)
von: Jing, Linglin, et al.
Veröffentlicht: (2025)
A Survey on Mixture of Experts in Large Language Models
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
von: Dong, Zican, et al.
Veröffentlicht: (2025)
von: Dong, Zican, et al.
Veröffentlicht: (2025)
Bayesian Mixture of Experts For Large Language Models
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
von: Dialameh, Maryam, et al.
Veröffentlicht: (2025)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
von: Xie, Yanyue, et al.
Veröffentlicht: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
von: Tan, Zheyue, et al.
Veröffentlicht: (2025)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
A Closer Look into Mixture-of-Experts in Large Language Models
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
von: Lo, Ka Man, et al.
Veröffentlicht: (2024)
Pruning General Large Language Models into Customized Expert Models
von: Zhao, Yirao, et al.
Veröffentlicht: (2025)
von: Zhao, Yirao, et al.
Veröffentlicht: (2025)
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
von: Dai, Damai, et al.
Veröffentlicht: (2024)
von: Dai, Damai, et al.
Veröffentlicht: (2024)
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
von: Li, Houyi, et al.
Veröffentlicht: (2025)
von: Li, Houyi, et al.
Veröffentlicht: (2025)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
von: Yoon, Youngsik, et al.
Veröffentlicht: (2026)
von: Yoon, Youngsik, et al.
Veröffentlicht: (2026)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
von: Gao, Shangqian, et al.
Veröffentlicht: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
MoLAE: Mixture of Latent Experts for Parameter-Efficient Language Models
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
Routing-Free Mixture-of-Experts
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
von: Liu, Yilun, et al.
Veröffentlicht: (2026)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
Layerwise Recurrent Router for Mixture-of-Experts
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
von: Qiu, Zihan, et al.
Veröffentlicht: (2024)
Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
von: Su, Guinan, et al.
Veröffentlicht: (2025)
von: Su, Guinan, et al.
Veröffentlicht: (2025)
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
von: Cai, Weilin, et al.
Veröffentlicht: (2024)
Mixture of Neuron Experts
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
von: Cheng, Runxi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
von: Guo, Hongcheng, et al.
Veröffentlicht: (2025) -
SPP: Sparsity-Preserved Parameter-Efficient Fine-Tuning for Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024) -
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
von: Huang, Yushi, et al.
Veröffentlicht: (2025) -
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024) -
Unveiling Super Experts in Mixture-of-Experts Large Language Models
von: Su, Zunhai, et al.
Veröffentlicht: (2025)