Unveiling Super Experts in Mixture-of-Experts Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Su, Zunhai, Li, Qingyuan, Zhang, Hao, Ye, Weihao, Xue, Qibo, Qian, YuLei, Xie, Yuchen, Wong, Ngai, Yuan, Kehong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915792025550848
author Su, Zunhai
Li, Qingyuan
Zhang, Hao
Ye, Weihao
Xue, Qibo
Qian, YuLei
Xie, Yuchen
Wong, Ngai
Yuan, Kehong
author_facet Su, Zunhai
Li, Qingyuan
Zhang, Hao
Ye, Weihao
Xue, Qibo
Qian, YuLei
Xie, Yuchen
Wong, Ngai
Yuan, Kehong
contents In this study, we report, for the first time, the discovery and systematic investigation of a distinct subset of experts that play a pivotal role in the MoE LLMs' forward inference. These experts are prevalent in open-source MoE LLMs, and despite their extremely limited number, pruning them results in a substantial decline in model performance (e.g., prune just three out of 6,144 causes Qwen3-30B-A3B to generate repetitive and uninformative outputs).We refer to these experts as Super Experts (SEs). Our comprehensive analysis provides progressively deeper insights into SEs: (i) SEs are characterized by rare but extreme activation outliers in the output of the down_proj, which give rise to massive activations in the hidden states between decoder layers. Moreover, the distribution of SEs is model-specific, data-agnostic, and remains unaffected by post-training processes. (ii) By pruning SEs, we assess their significance across a variety of tasks, revealing their considerable impact on the model's overall performance, particularly in mathematical reasoning. (iii) We further investigate why compressing SEs exerts such a pronounced impact. We show that, in MoE LLMs, SEs serve as the primary source of the systematic outlier mechanism in Transformers, and that compressing them profoundly disrupts this process, ultimately causing the collapse of attention sinks. These findings advance the understanding of the internal dynamics of MoE LLMs, filling an important gap in the current knowledge. The code is provided in https://github.com/ZunhaiSu/Super-Experts-Profilling.
format Preprint
id arxiv_https___arxiv_org_abs_2507_23279
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unveiling Super Experts in Mixture-of-Experts Large Language Models
Su, Zunhai
Li, Qingyuan
Zhang, Hao
Ye, Weihao
Xue, Qibo
Qian, YuLei
Xie, Yuchen
Wong, Ngai
Yuan, Kehong
Computation and Language
In this study, we report, for the first time, the discovery and systematic investigation of a distinct subset of experts that play a pivotal role in the MoE LLMs' forward inference. These experts are prevalent in open-source MoE LLMs, and despite their extremely limited number, pruning them results in a substantial decline in model performance (e.g., prune just three out of 6,144 causes Qwen3-30B-A3B to generate repetitive and uninformative outputs).We refer to these experts as Super Experts (SEs). Our comprehensive analysis provides progressively deeper insights into SEs: (i) SEs are characterized by rare but extreme activation outliers in the output of the down_proj, which give rise to massive activations in the hidden states between decoder layers. Moreover, the distribution of SEs is model-specific, data-agnostic, and remains unaffected by post-training processes. (ii) By pruning SEs, we assess their significance across a variety of tasks, revealing their considerable impact on the model's overall performance, particularly in mathematical reasoning. (iii) We further investigate why compressing SEs exerts such a pronounced impact. We show that, in MoE LLMs, SEs serve as the primary source of the systematic outlier mechanism in Transformers, and that compressing them profoundly disrupts this process, ultimately causing the collapse of attention sinks. These findings advance the understanding of the internal dynamics of MoE LLMs, filling an important gap in the current knowledge. The code is provided in https://github.com/ZunhaiSu/Super-Experts-Profilling.
title Unveiling Super Experts in Mixture-of-Experts Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.23279