SD-MoE: Spectral Decomposition for Effective Expert Specialization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Ruijun, Dong, Fang, Zhang, Xin, Cao, Hengjie, Huang, Zhendong, Chen, Anrui, Zhou, Jixian, Chen, Mengyi, Yang, Yifeng, Dong, Mingzhi, Wang, Yujiang, Hou, Jinlong, Lv, Qin, Dick, Robert P., Cheng, Yuan, Yang, Fan, Lu, Tun, Zhang, Chun, Shang, Li
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914327768858624
author Huang, Ruijun
Dong, Fang
Zhang, Xin
Cao, Hengjie
Huang, Zhendong
Chen, Anrui
Zhou, Jixian
Chen, Mengyi
Yang, Yifeng
Dong, Mingzhi
Wang, Yujiang
Hou, Jinlong
Lv, Qin
Dick, Robert P.
Cheng, Yuan
Yang, Fan
Lu, Tun
Zhang, Chun
Shang, Li
author_facet Huang, Ruijun
Dong, Fang
Zhang, Xin
Cao, Hengjie
Huang, Zhendong
Chen, Anrui
Zhou, Jixian
Chen, Mengyi
Yang, Yifeng
Dong, Mingzhi
Wang, Yujiang
Hou, Jinlong
Lv, Qin
Dick, Robert P.
Cheng, Yuan
Yang, Fan
Lu, Tun
Zhang, Chun
Shang, Li
contents Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often fails: some experts become functionally similar, while others functioning as de facto shared experts, limiting the effective capacity and model performance. In this work, we analysis from a spectral perspective on parameter and gradient spaces, uncover that (1) experts share highly overlapping dominant spectral components in their parameters, (2) dominant gradient subspaces are strongly aligned across experts, driven by ubiquitous low-rank structure in human corpus, and (3) gating mechanisms preferentially route inputs along these dominant directions, further limiting specialization. To address this, we propose Spectral-Decoupled MoE (SD-MoE), which decomposes both parameter and gradient in the spectral space. SD-MoE improves performance across downstream tasks, enables effective expert specialization, incurring minimal additional computation, and can be seamlessly integrated into a wide range of existing MoE architectures, including Qwen and DeepSeek.
format Preprint
id arxiv_https___arxiv_org_abs_2602_12556
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SD-MoE: Spectral Decomposition for Effective Expert Specialization
Huang, Ruijun
Dong, Fang
Zhang, Xin
Cao, Hengjie
Huang, Zhendong
Chen, Anrui
Zhou, Jixian
Chen, Mengyi
Yang, Yifeng
Dong, Mingzhi
Wang, Yujiang
Hou, Jinlong
Lv, Qin
Dick, Robert P.
Cheng, Yuan
Yang, Fan
Lu, Tun
Zhang, Chun
Shang, Li
Machine Learning
Artificial Intelligence
Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often fails: some experts become functionally similar, while others functioning as de facto shared experts, limiting the effective capacity and model performance. In this work, we analysis from a spectral perspective on parameter and gradient spaces, uncover that (1) experts share highly overlapping dominant spectral components in their parameters, (2) dominant gradient subspaces are strongly aligned across experts, driven by ubiquitous low-rank structure in human corpus, and (3) gating mechanisms preferentially route inputs along these dominant directions, further limiting specialization. To address this, we propose Spectral-Decoupled MoE (SD-MoE), which decomposes both parameter and gradient in the spectral space. SD-MoE improves performance across downstream tasks, enables effective expert specialization, incurring minimal additional computation, and can be seamlessly integrated into a wide range of existing MoE architectures, including Qwen and DeepSeek.
title SD-MoE: Spectral Decomposition for Effective Expert Specialization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.12556