DBES: A Systematic Benchmark and Metric Suite for Evaluating Expert Specialization in Large-Scale MoEs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Jing, Lu, Hongxuan, Young, Jazze, Wang, Shu, Xin, Zhimin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SD-MoE: Spectral Decomposition for Effective Expert Specialization
par: Huang, Ruijun, et autres
Publié: (2026)
par: Huang, Ruijun, et autres
Publié: (2026)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
par: Jaiswal, Ajay, et autres
Publié: (2025)
par: Jaiswal, Ajay, et autres
Publié: (2025)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
par: Farhat, Yehya, et autres
Publié: (2023)
par: Farhat, Yehya, et autres
Publié: (2023)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
par: Ye, Charles, et autres
Publié: (2026)
par: Ye, Charles, et autres
Publié: (2026)
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
par: Wang, Xi, et autres
Publié: (2026)
par: Wang, Xi, et autres
Publié: (2026)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
par: Chen, Yuanteng, et autres
Publié: (2025)
par: Chen, Yuanteng, et autres
Publié: (2025)
Collaborative Compression for Large-Scale MoE Deployment on Edge
par: Chen, Yixiao, et autres
Publié: (2025)
par: Chen, Yixiao, et autres
Publié: (2025)
Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation
par: Li, Junzhuo, et autres
Publié: (2025)
par: Li, Junzhuo, et autres
Publié: (2025)
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
par: Shi, Xiaoming, et autres
Publié: (2024)
par: Shi, Xiaoming, et autres
Publié: (2024)
Expert Divergence Learning for MoE-based Language Models
par: Li, Jiaang, et autres
Publié: (2026)
par: Li, Jiaang, et autres
Publié: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
LoRALib: A Standardized Benchmark for Evaluating LoRA-MoE Methods
par: Wang, Shaoheng, et autres
Publié: (2025)
par: Wang, Shaoheng, et autres
Publié: (2025)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
par: Jin, Peng, et autres
Publié: (2024)
par: Jin, Peng, et autres
Publié: (2024)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
par: Neogi, Pinaki Prasad Guha, et autres
Publié: (2025)
par: Neogi, Pinaki Prasad Guha, et autres
Publié: (2025)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
par: Ye, Xin, et autres
Publié: (2026)
par: Ye, Xin, et autres
Publié: (2026)
MoE-Health: A Mixture of Experts Framework for Robust Multimodal Healthcare Prediction
par: Wang, Xiaoyang, et autres
Publié: (2025)
par: Wang, Xiaoyang, et autres
Publié: (2025)
Practical FP4 Training for Large-Scale MoE Models on Hopper GPUs
par: Zhang, Wuyue, et autres
Publié: (2026)
par: Zhang, Wuyue, et autres
Publié: (2026)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
par: Ludziejewski, Jan, et autres
Publié: (2025)
par: Ludziejewski, Jan, et autres
Publié: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
par: Gu, Naibin, et autres
Publié: (2025)
par: Gu, Naibin, et autres
Publié: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Mixture of Experts (MoE): A Big Data Perspective
par: Gan, Wensheng, et autres
Publié: (2025)
par: Gan, Wensheng, et autres
Publié: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
par: Kulibaba, Stepan, et autres
Publié: (2026)
par: Kulibaba, Stepan, et autres
Publié: (2026)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
par: Zibakhsh, Soheil, et autres
Publié: (2025)
par: Zibakhsh, Soheil, et autres
Publié: (2025)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
par: Wang, Qi, et autres
Publié: (2025)
par: Wang, Qi, et autres
Publié: (2025)
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
par: Yun, Sukwon, et autres
Publié: (2024)
par: Yun, Sukwon, et autres
Publié: (2024)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
par: Wu, Haoze, et autres
Publié: (2024)
par: Wu, Haoze, et autres
Publié: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
par: Liu, Baihui, et autres
Publié: (2026)
par: Liu, Baihui, et autres
Publié: (2026)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
par: Su, Yang, et autres
Publié: (2025)
par: Su, Yang, et autres
Publié: (2025)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
par: Hu, Wentao, et autres
Publié: (2026)
par: Hu, Wentao, et autres
Publié: (2026)
FFT-MoE: Efficient Federated Fine-Tuning for Foundation Models via Large-scale Sparse MoE under Heterogeneous Edge
par: Hu, Gang, et autres
Publié: (2025)
par: Hu, Gang, et autres
Publié: (2025)
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
par: Wang, Siqi, et autres
Publié: (2024)
par: Wang, Siqi, et autres
Publié: (2024)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
par: Li, Pingzhi, et autres
Publié: (2025)
par: Li, Pingzhi, et autres
Publié: (2025)
Continual Pre-training of MoEs: How robust is your router?
par: Thérien, Benjamin, et autres
Publié: (2025)
par: Thérien, Benjamin, et autres
Publié: (2025)
LocMoE: A Low-Overhead MoE for Large Language Model Training
par: Li, Jing, et autres
Publié: (2024)
par: Li, Jing, et autres
Publié: (2024)
BrainNet-MoE: Brain-Inspired Mixture-of-Experts Learning for Neurological Disease Identification
par: Zhang, Jing, et autres
Publié: (2025)
par: Zhang, Jing, et autres
Publié: (2025)
MP-MoE: Matrix Profile-Guided Mixture of Experts for Precipitation Forecasting
par: Tran, Huyen Ngoc, et autres
Publié: (2026)
par: Tran, Huyen Ngoc, et autres
Publié: (2026)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
par: Xu, Zukang, et autres
Publié: (2026)
par: Xu, Zukang, et autres
Publié: (2026)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
par: Hannah, Lauren. A, et autres
Publié: (2025)
par: Hannah, Lauren. A, et autres
Publié: (2025)
Accelerating MoE Model Inference with Expert Sharding
par: Balmau, Oana, et autres
Publié: (2025)
par: Balmau, Oana, et autres
Publié: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
par: Yang, Cheng, et autres
Publié: (2024)
par: Yang, Cheng, et autres
Publié: (2024)
Documents similaires
-
SD-MoE: Spectral Decomposition for Effective Expert Specialization
par: Huang, Ruijun, et autres
Publié: (2026) -
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
par: Jaiswal, Ajay, et autres
Publié: (2025) -
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
par: Farhat, Yehya, et autres
Publié: (2023) -
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
par: Ye, Charles, et autres
Publié: (2026) -
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
par: Wang, Xi, et autres
Publié: (2026)