PreMoE: Proactive Inference for Efficient Mixture-of-Experts
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911620200923136 |
|---|---|
| author | Pei, Zehua Zhang, Ying Zhen, Hui-Ling Yuan, Tao Yu, Xianzhi Dong, Zhenhua Pan, Sinno Jialin Yuan, Mingxuan Yu, Bei |
| author_facet | Pei, Zehua Zhang, Ying Zhen, Hui-Ling Yuan, Tao Yu, Xianzhi Dong, Zhenhua Pan, Sinno Jialin Yuan, Mingxuan Yu, Bei |
| contents | Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization. We introduce PreMoE, a training-free framework that proactively compiles sparse MoE variants for targeted deployment scenarios. At its core is Predicted Expert Utility (PEU), a robust metric for estimating expert importance from router logits through high-confidence threshold filtering and logit transformation, which together stabilize utility estimation under aggressive sparsity. Using PEU scores computed on a small calibration set, PreMoE produces domain-aware expert rankings that can be used to compile either domain-specific specialists or high-efficiency multi-domain generalists, without any retraining. Across MoE models ranging from 30B to 718B parameters, PreMoE achieves up to 50\% sparsity with nearly no performance loss. It further exposes a practical deployment trade-off: specialists maximize in-domain efficiency, while synthesized generalists retain broader cross-domain capability at the same sparsity budget. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_17639 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | PreMoE: Proactive Inference for Efficient Mixture-of-Experts Pei, Zehua Zhang, Ying Zhen, Hui-Ling Yuan, Tao Yu, Xianzhi Dong, Zhenhua Pan, Sinno Jialin Yuan, Mingxuan Yu, Bei Machine Learning Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization. We introduce PreMoE, a training-free framework that proactively compiles sparse MoE variants for targeted deployment scenarios. At its core is Predicted Expert Utility (PEU), a robust metric for estimating expert importance from router logits through high-confidence threshold filtering and logit transformation, which together stabilize utility estimation under aggressive sparsity. Using PEU scores computed on a small calibration set, PreMoE produces domain-aware expert rankings that can be used to compile either domain-specific specialists or high-efficiency multi-domain generalists, without any retraining. Across MoE models ranging from 30B to 718B parameters, PreMoE achieves up to 50\% sparsity with nearly no performance loss. It further exposes a practical deployment trade-off: specialists maximize in-domain efficiency, while synthesized generalists retain broader cross-domain capability at the same sparsity budget. |
| title | PreMoE: Proactive Inference for Efficient Mixture-of-Experts |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.17639 |