Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hyeon, Sieun, Do, Jaeyoung
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917306173489152
author Hyeon, Sieun
Do, Jaeyoung
author_facet Hyeon, Sieun
Do, Jaeyoung
contents Mixture-of-Experts (MoE) models scale capacity efficiently, but their massive parameter footprint creates a deployment-time memory bottleneck. We organize retraining-free MoE compression into three paradigms - Expert Pruning, Expert Editing, and Expert Merging - and show that persistent post-compression degradation largely stems from a neglected factor: router-expert mismatch when experts are changed but the router is left untouched. We argue that effective retraining-free compression should avoid updating expert parameters while allowing lightweight router calibration. To this end, we propose Router Knowledge Distillation (Router KD), which updates only a tiny fraction of parameters (the router) by distilling the original model's next-token distribution on unlabeled calibration data. Experiments across representative methods in all three paradigms demonstrate consistent performance recovery, with substantially larger gains in fine-grained MoEs (many small experts) than in coarse-grained MoEs due to their more complex routing decision boundaries.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02217
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
Hyeon, Sieun
Do, Jaeyoung
Machine Learning
Artificial Intelligence
Mixture-of-Experts (MoE) models scale capacity efficiently, but their massive parameter footprint creates a deployment-time memory bottleneck. We organize retraining-free MoE compression into three paradigms - Expert Pruning, Expert Editing, and Expert Merging - and show that persistent post-compression degradation largely stems from a neglected factor: router-expert mismatch when experts are changed but the router is left untouched. We argue that effective retraining-free compression should avoid updating expert parameters while allowing lightweight router calibration. To this end, we propose Router Knowledge Distillation (Router KD), which updates only a tiny fraction of parameters (the router) by distilling the original model's next-token distribution on unlabeled calibration data. Experiments across representative methods in all three paradigms demonstrate consistent performance recovery, with substantially larger gains in fine-grained MoEs (many small experts) than in coarse-grained MoEs due to their more complex routing decision boundaries.
title Is Retraining-Free Enough? The Necessity of Router Calibration for Efficient MoE Compression
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.02217