Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chu, Sanghyeok, Ahn, Pyunghwan, Song, Gwangmo, Kim, SeungHwan, Lee, Honglak, Han, Bohyung
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915941636374528
author Chu, Sanghyeok
Ahn, Pyunghwan
Song, Gwangmo
Kim, SeungHwan
Lee, Honglak
Han, Bohyung
author_facet Chu, Sanghyeok
Ahn, Pyunghwan
Song, Gwangmo
Kim, SeungHwan
Lee, Honglak
Han, Bohyung
contents Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch. However, since all experts start from identical weights and the router is randomly initialized, the model suffers from expert symmetry and limited early specialization. We propose Cluster-aware Upcycling, a strategy that incorporates semantic structure into MoE initialization. Our method first partitions the dense model's input activations into semantic clusters. Each expert is then initialized using the subspace representations of its corresponding cluster via truncated SVD, while setting the router's initial weights to the cluster centroids. This cluster-aware initialization breaks expert symmetry and encourages early specialization aligned with the data distribution. Furthermore, we introduce an expert-ensemble self-distillation loss that stabilizes training by providing reliable routing guidance using an ensemble teacher. When evaluated on CLIP ViT-B/32 and ViT-B/16, Cluster-aware Upcycling consistently outperforms existing methods across both zero-shot and few-shot benchmarks. The proposed method also produces more diverse and disentangled expert representations, reduces inter-expert similarity, and leads to more confident routing behavior. Project page: https://sanghyeokchu.github.io/cluster-aware-upcycling/
format Preprint
id arxiv_https___arxiv_org_abs_2604_13508
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
Chu, Sanghyeok
Ahn, Pyunghwan
Song, Gwangmo
Kim, SeungHwan
Lee, Honglak
Han, Bohyung
Computer Vision and Pattern Recognition
Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch. However, since all experts start from identical weights and the router is randomly initialized, the model suffers from expert symmetry and limited early specialization. We propose Cluster-aware Upcycling, a strategy that incorporates semantic structure into MoE initialization. Our method first partitions the dense model's input activations into semantic clusters. Each expert is then initialized using the subspace representations of its corresponding cluster via truncated SVD, while setting the router's initial weights to the cluster centroids. This cluster-aware initialization breaks expert symmetry and encourages early specialization aligned with the data distribution. Furthermore, we introduce an expert-ensemble self-distillation loss that stabilizes training by providing reliable routing guidance using an ensemble teacher. When evaluated on CLIP ViT-B/32 and ViT-B/16, Cluster-aware Upcycling consistently outperforms existing methods across both zero-shot and few-shot benchmarks. The proposed method also produces more diverse and disentangled expert representations, reduces inter-expert similarity, and leads to more confident routing behavior. Project page: https://sanghyeokchu.github.io/cluster-aware-upcycling/
title Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.13508