The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xi, Hayou, Soufiane, Nalisnick, Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Proof of Learning Rate Transfer under $μ$P
von: Hayou, Soufiane
Veröffentlicht: (2025)
von: Hayou, Soufiane
Veröffentlicht: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
von: Hayou, Soufiane, et al.
Veröffentlicht: (2025)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2025)
DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
von: Wang, Jing, et al.
Veröffentlicht: (2026)
von: Wang, Jing, et al.
Veröffentlicht: (2026)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
Learning Rate Scaling across LoRA Ranks and Transfer to Full Finetuning
von: Chen, Nan, et al.
Veröffentlicht: (2026)
von: Chen, Nan, et al.
Veröffentlicht: (2026)
The Impact of Initialization on LoRA Finetuning Dynamics
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
LoRA+: Efficient Low Rank Adaptation of Large Models
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
von: Hayou, Soufiane, et al.
Veröffentlicht: (2024)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2025)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2025)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
von: Li, Junzhuo, et al.
Veröffentlicht: (2025)
MoE-MLoRA for Multi-Domain CTR Prediction: Efficient Adaptation with Expert Specialization
von: Yaggel, Ken, et al.
Veröffentlicht: (2025)
von: Yaggel, Ken, et al.
Veröffentlicht: (2025)
Are vision language models robust to uncertain inputs?
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
BrainStack: Neuro-MoE with Functionally Guided Expert Routing for EEG-Based Language Decoding
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziyi, et al.
Veröffentlicht: (2026)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
von: Lasby, Mike, et al.
Veröffentlicht: (2025)
von: Lasby, Mike, et al.
Veröffentlicht: (2025)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
MESA: Improving MoE Safety Alignment via Decentralized Expertise
von: Sun, Yitong, et al.
Veröffentlicht: (2026)
von: Sun, Yitong, et al.
Veröffentlicht: (2026)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
MoE-Loco: Mixture of Experts for Multitask Locomotion
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
von: Wang, Haodong, et al.
Veröffentlicht: (2025)
Expert Divergence Learning for MoE-based Language Models
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
von: Li, Jiaang, et al.
Veröffentlicht: (2026)
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering
von: Zeng, Xiyin, et al.
Veröffentlicht: (2026)
von: Zeng, Xiyin, et al.
Veröffentlicht: (2026)
$μ$pscaling small models: Principled warm starts and hyperparameter transfer
von: Ma, Yuxin, et al.
Veröffentlicht: (2026)
von: Ma, Yuxin, et al.
Veröffentlicht: (2026)
MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
von: Shi, Jingze, et al.
Veröffentlicht: (2026)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
von: Feng, Jiarui, et al.
Veröffentlicht: (2026)
von: Feng, Jiarui, et al.
Veröffentlicht: (2026)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
von: Shen, Weijie, et al.
Veröffentlicht: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
von: Kulibaba, Stepan, et al.
Veröffentlicht: (2026)
Mixture of Experts (MoE): A Big Data Perspective
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
von: Gan, Wensheng, et al.
Veröffentlicht: (2025)
Unveiling and Consulting Core Experts in Retrieval-Augmented MoE-based LLMs
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
von: Zhou, Xin, et al.
Veröffentlicht: (2024)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
von: Gu, Naibin, et al.
Veröffentlicht: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
Continual Pre-training of MoEs: How robust is your router?
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Proof of Learning Rate Transfer under $μ$P
von: Hayou, Soufiane
Veröffentlicht: (2025) -
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
von: Ye, Charles, et al.
Veröffentlicht: (2026) -
Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size
von: Hayou, Soufiane, et al.
Veröffentlicht: (2025) -
DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
von: Wang, Jing, et al.
Veröffentlicht: (2026) -
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)