PreMoE: Proactive Inference for Efficient Mixture-of-Experts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pei, Zehua, Zhang, Ying, Zhen, Hui-Ling, Yuan, Tao, Yu, Xianzhi, Dong, Zhenhua, Pan, Sinno Jialin, Yuan, Mingxuan, Yu, Bei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911620200923136
author Pei, Zehua
Zhang, Ying
Zhen, Hui-Ling
Yuan, Tao
Yu, Xianzhi
Dong, Zhenhua
Pan, Sinno Jialin
Yuan, Mingxuan
Yu, Bei
author_facet Pei, Zehua
Zhang, Ying
Zhen, Hui-Ling
Yuan, Tao
Yu, Xianzhi
Dong, Zhenhua
Pan, Sinno Jialin
Yuan, Mingxuan
Yu, Bei
contents Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization. We introduce PreMoE, a training-free framework that proactively compiles sparse MoE variants for targeted deployment scenarios. At its core is Predicted Expert Utility (PEU), a robust metric for estimating expert importance from router logits through high-confidence threshold filtering and logit transformation, which together stabilize utility estimation under aggressive sparsity. Using PEU scores computed on a small calibration set, PreMoE produces domain-aware expert rankings that can be used to compile either domain-specific specialists or high-efficiency multi-domain generalists, without any retraining. Across MoE models ranging from 30B to 718B parameters, PreMoE achieves up to 50\% sparsity with nearly no performance loss. It further exposes a practical deployment trade-off: specialists maximize in-domain efficiency, while synthesized generalists retain broader cross-domain capability at the same sparsity budget.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17639
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PreMoE: Proactive Inference for Efficient Mixture-of-Experts
Pei, Zehua
Zhang, Ying
Zhen, Hui-Ling
Yuan, Tao
Yu, Xianzhi
Dong, Zhenhua
Pan, Sinno Jialin
Yuan, Mingxuan
Yu, Bei
Machine Learning
Mixture-of-Experts (MoE) models offer dynamic computation, but are typically deployed as static full-capacity models, missing opportunities for deployment-specific specialization. We introduce PreMoE, a training-free framework that proactively compiles sparse MoE variants for targeted deployment scenarios. At its core is Predicted Expert Utility (PEU), a robust metric for estimating expert importance from router logits through high-confidence threshold filtering and logit transformation, which together stabilize utility estimation under aggressive sparsity. Using PEU scores computed on a small calibration set, PreMoE produces domain-aware expert rankings that can be used to compile either domain-specific specialists or high-efficiency multi-domain generalists, without any retraining. Across MoE models ranging from 30B to 718B parameters, PreMoE achieves up to 50\% sparsity with nearly no performance loss. It further exposes a practical deployment trade-off: specialists maximize in-domain efficiency, while synthesized generalists retain broader cross-domain capability at the same sparsity budget.
title PreMoE: Proactive Inference for Efficient Mixture-of-Experts
topic Machine Learning
url https://arxiv.org/abs/2505.17639