ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Anzhe, Duan, Shukai, Li, Shixuan, Yin, Chenzhong, Cheng, Mingxi, Ping, Heng, Chattopadhyay, Tamoghna, Thomopoulos, Sophia I, Nazarian, Shahin, Thompson, Paul, Bogdan, Paul
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918411374690304
author Cheng, Anzhe
Duan, Shukai
Li, Shixuan
Yin, Chenzhong
Cheng, Mingxi
Ping, Heng
Chattopadhyay, Tamoghna
Thomopoulos, Sophia I
Nazarian, Shahin
Thompson, Paul
Bogdan, Paul
author_facet Cheng, Anzhe
Duan, Shukai
Li, Shixuan
Yin, Chenzhong
Cheng, Mingxi
Ping, Heng
Chattopadhyay, Tamoghna
Thomopoulos, Sophia I
Nazarian, Shahin
Thompson, Paul
Bogdan, Paul
contents Mixture-of-Experts (MoE) architectures expand model capacity by sparsely activating experts but face two core challenges: misalignment between router logits and each expert's internal structure leads to unstable routing and expert underutilization, and load imbalances create straggler bottlenecks. Standard solutions, such as auxiliary load-balancing losses, can reduce load disparities but often weaken expert specialization and hurt downstream performance. To address these issues, we propose ERMoE, a sparse MoE transformer that reparameterizes each expert in a learned orthonormal eigenbasis and replaces learned gating logits with an "Eigenbasis Score", defined as the cosine similarity between input features and an expert's basis. This content-aware routing ties token assignments directly to experts' representation spaces, stabilizing utilization and promoting interpretable specialization without sacrificing sparsity. Crucially, ERMoE removes the need for explicit balancing losses and avoids the interfering gradients they introduce. We show that ERMoE achieves state-of-the-art accuracy on ImageNet classification and cross-modal image-text retrieval benchmarks (e.g., COCO, Flickr30K), while naturally producing flatter expert load distributions. Moreover, a 3D MRI variant (ERMoE-ba) improves brain age prediction accuracy by more than 7\% and yields anatomically interpretable expert specializations. ERMoE thus introduces a new architectural principle for sparse expert models that directly addresses routing instabilities and enables improved performance with scalable, interpretable specialization.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10971
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
Cheng, Anzhe
Duan, Shukai
Li, Shixuan
Yin, Chenzhong
Cheng, Mingxi
Ping, Heng
Chattopadhyay, Tamoghna
Thomopoulos, Sophia I
Nazarian, Shahin
Thompson, Paul
Bogdan, Paul
Computer Vision and Pattern Recognition
Mixture-of-Experts (MoE) architectures expand model capacity by sparsely activating experts but face two core challenges: misalignment between router logits and each expert's internal structure leads to unstable routing and expert underutilization, and load imbalances create straggler bottlenecks. Standard solutions, such as auxiliary load-balancing losses, can reduce load disparities but often weaken expert specialization and hurt downstream performance. To address these issues, we propose ERMoE, a sparse MoE transformer that reparameterizes each expert in a learned orthonormal eigenbasis and replaces learned gating logits with an "Eigenbasis Score", defined as the cosine similarity between input features and an expert's basis. This content-aware routing ties token assignments directly to experts' representation spaces, stabilizing utilization and promoting interpretable specialization without sacrificing sparsity. Crucially, ERMoE removes the need for explicit balancing losses and avoids the interfering gradients they introduce. We show that ERMoE achieves state-of-the-art accuracy on ImageNet classification and cross-modal image-text retrieval benchmarks (e.g., COCO, Flickr30K), while naturally producing flatter expert load distributions. Moreover, a 3D MRI variant (ERMoE-ba) improves brain age prediction accuracy by more than 7\% and yields anatomically interpretable expert specializations. ERMoE thus introduces a new architectural principle for sparse expert models that directly addresses routing instabilities and enables improved performance with scalable, interpretable specialization.
title ERMoE: Eigen-Reparameterized Mixture-of-Experts for Stable Routing and Interpretable Specialization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.10971