SD-MoE: Spectral Decomposition for Effective Expert Specialization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914327768858624 |
|---|---|
| author | Huang, Ruijun Dong, Fang Zhang, Xin Cao, Hengjie Huang, Zhendong Chen, Anrui Zhou, Jixian Chen, Mengyi Yang, Yifeng Dong, Mingzhi Wang, Yujiang Hou, Jinlong Lv, Qin Dick, Robert P. Cheng, Yuan Yang, Fan Lu, Tun Zhang, Chun Shang, Li |
| author_facet | Huang, Ruijun Dong, Fang Zhang, Xin Cao, Hengjie Huang, Zhendong Chen, Anrui Zhou, Jixian Chen, Mengyi Yang, Yifeng Dong, Mingzhi Wang, Yujiang Hou, Jinlong Lv, Qin Dick, Robert P. Cheng, Yuan Yang, Fan Lu, Tun Zhang, Chun Shang, Li |
| contents | Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often fails: some experts become functionally similar, while others functioning as de facto shared experts, limiting the effective capacity and model performance. In this work, we analysis from a spectral perspective on parameter and gradient spaces, uncover that (1) experts share highly overlapping dominant spectral components in their parameters, (2) dominant gradient subspaces are strongly aligned across experts, driven by ubiquitous low-rank structure in human corpus, and (3) gating mechanisms preferentially route inputs along these dominant directions, further limiting specialization. To address this, we propose Spectral-Decoupled MoE (SD-MoE), which decomposes both parameter and gradient in the spectral space. SD-MoE improves performance across downstream tasks, enables effective expert specialization, incurring minimal additional computation, and can be seamlessly integrated into a wide range of existing MoE architectures, including Qwen and DeepSeek. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_12556 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | SD-MoE: Spectral Decomposition for Effective Expert Specialization Huang, Ruijun Dong, Fang Zhang, Xin Cao, Hengjie Huang, Zhendong Chen, Anrui Zhou, Jixian Chen, Mengyi Yang, Yifeng Dong, Mingzhi Wang, Yujiang Hou, Jinlong Lv, Qin Dick, Robert P. Cheng, Yuan Yang, Fan Lu, Tun Zhang, Chun Shang, Li Machine Learning Artificial Intelligence Mixture-of-Experts (MoE) architectures scale Large Language Models via expert specialization induced by conditional computation. In practice, however, expert specialization often fails: some experts become functionally similar, while others functioning as de facto shared experts, limiting the effective capacity and model performance. In this work, we analysis from a spectral perspective on parameter and gradient spaces, uncover that (1) experts share highly overlapping dominant spectral components in their parameters, (2) dominant gradient subspaces are strongly aligned across experts, driven by ubiquitous low-rank structure in human corpus, and (3) gating mechanisms preferentially route inputs along these dominant directions, further limiting specialization. To address this, we propose Spectral-Decoupled MoE (SD-MoE), which decomposes both parameter and gradient in the spectral space. SD-MoE improves performance across downstream tasks, enables effective expert specialization, incurring minimal additional computation, and can be seamlessly integrated into a wide range of existing MoE architectures, including Qwen and DeepSeek. |
| title | SD-MoE: Spectral Decomposition for Effective Expert Specialization |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.12556 |