How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Park, Sumin, Park, Noseong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
por: Choi, Jeongwhan, et al.
Publicado: (2025)
por: Choi, Jeongwhan, et al.
Publicado: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024)
por: Zhuang, Haomin, et al.
Publicado: (2024)
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
por: Choi, Jeongwhan, et al.
Publicado: (2024)
por: Choi, Jeongwhan, et al.
Publicado: (2024)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
por: Nikolic, Strahinja, et al.
Publicado: (2025)
por: Nikolic, Strahinja, et al.
Publicado: (2025)
Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
por: Choi, Hahyeon, et al.
Publicado: (2026)
por: Choi, Hahyeon, et al.
Publicado: (2026)
SPI-GAN: Denoising Diffusion GANs with Straight-Path Interpolations
por: Jeon, Jinsung, et al.
Publicado: (2022)
por: Jeon, Jinsung, et al.
Publicado: (2022)
HINTS: Extraction of Human Insights from Time-Series Without External Sources
por: Jhin, Sheo Yon, et al.
Publicado: (2025)
por: Jhin, Sheo Yon, et al.
Publicado: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
por: Madan, Vivan, et al.
Publicado: (2026)
por: Madan, Vivan, et al.
Publicado: (2026)
Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design
por: Li, Junzhuo, et al.
Publicado: (2026)
por: Li, Junzhuo, et al.
Publicado: (2026)
LatentMoE: Toward Optimal Accuracy per FLOP and Parameter in Mixture of Experts
por: Elango, Venmugil, et al.
Publicado: (2026)
por: Elango, Venmugil, et al.
Publicado: (2026)
HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts
por: Zhao, Hao, et al.
Publicado: (2024)
por: Zhao, Hao, et al.
Publicado: (2024)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
por: Park, Jongseok, et al.
Publicado: (2026)
por: Park, Jongseok, et al.
Publicado: (2026)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
por: Bambhaniya, Abhimanyu, et al.
Publicado: (2026)
Mixture of Raytraced Experts
por: Perin, Andrea, et al.
Publicado: (2025)
por: Perin, Andrea, et al.
Publicado: (2025)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
por: Gao, Yuting, et al.
Publicado: (2025)
por: Gao, Yuting, et al.
Publicado: (2025)
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
por: He, Yifei, et al.
Publicado: (2025)
por: He, Yifei, et al.
Publicado: (2025)
MoDEx: Mixture of Depth-specific Experts for Multivariate Long-term Time Series Forecasting
por: Yoon, Hyekyung, et al.
Publicado: (2026)
por: Yoon, Hyekyung, et al.
Publicado: (2026)
The Illusion of Specialization: Unveiling the Domain-Invariant "Standing Committee" in Mixture-of-Experts Models
por: Wang, Yan, et al.
Publicado: (2026)
por: Wang, Yan, et al.
Publicado: (2026)
One Prompt is not Enough: Automated Construction of a Mixture-of-Expert Prompts
por: Wang, Ruochen, et al.
Publicado: (2024)
por: Wang, Ruochen, et al.
Publicado: (2024)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
por: Wi, Hyowon, et al.
Publicado: (2025)
por: Wi, Hyowon, et al.
Publicado: (2025)
Graph Signal Processing Meets Mamba2: Adaptive Filter Bank via Delta Modulation
por: Shin, Yehjin, et al.
Publicado: (2026)
por: Shin, Yehjin, et al.
Publicado: (2026)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
por: Gritsch, Nikolas, et al.
Publicado: (2024)
por: Gritsch, Nikolas, et al.
Publicado: (2024)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
por: Yan, Jiaming, et al.
Publicado: (2025)
por: Yan, Jiaming, et al.
Publicado: (2025)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
por: Dwivedi, Chaitanya, et al.
Publicado: (2026)
por: Dwivedi, Chaitanya, et al.
Publicado: (2026)
Sparsity and Superposition in Mixture of Experts
por: Chaudhari, Marmik, et al.
Publicado: (2025)
por: Chaudhari, Marmik, et al.
Publicado: (2025)
Mixture of Diverse Size Experts
por: Sun, Manxi, et al.
Publicado: (2024)
por: Sun, Manxi, et al.
Publicado: (2024)
Mixture of Concept Bottleneck Experts
por: De Santis, Francesco, et al.
Publicado: (2026)
por: De Santis, Francesco, et al.
Publicado: (2026)
Mixture of A Million Experts
por: He, Xu Owen
Publicado: (2024)
por: He, Xu Owen
Publicado: (2024)
Mixture of Experts in a Mixture of RL settings
por: Willi, Timon, et al.
Publicado: (2024)
por: Willi, Timon, et al.
Publicado: (2024)
Probing Semantic Routing in Large Mixture-of-Expert Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
por: Chu, Kexin, et al.
Publicado: (2025)
por: Chu, Kexin, et al.
Publicado: (2025)
Peirce in the Machine: How Mixture of Experts Models Perform Hypothesis Construction
por: Rushing, Bruce
Publicado: (2024)
por: Rushing, Bruce
Publicado: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
por: Nguyen-Nhat, Minh-Khoi, et al.
Publicado: (2025)
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts
por: Huang, Minbin, et al.
Publicado: (2026)
por: Huang, Minbin, et al.
Publicado: (2026)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
por: Yan, Fanqi, et al.
Publicado: (2024)
por: Yan, Fanqi, et al.
Publicado: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
por: Nguyen, Xuan-Phi, et al.
Publicado: (2026)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
por: Kim, Gyeongman, et al.
Publicado: (2025)
por: Kim, Gyeongman, et al.
Publicado: (2025)
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
por: He, Shwai, et al.
Publicado: (2024)
por: He, Shwai, et al.
Publicado: (2024)
Towards Generalization-Oriented Models for Vehicle Routing Problems with Mixture-of-Experts
por: Miao, Changhao, et al.
Publicado: (2026)
por: Miao, Changhao, et al.
Publicado: (2026)
Ejemplares similares
-
Are Graph Transformers Necessary? Efficient Long-Range Message Passing with Fractal Nodes in MPNNs
por: Choi, Jeongwhan, et al.
Publicado: (2025) -
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024) -
PANDA: Expanded Width-Aware Message Passing Beyond Rewiring
por: Choi, Jeongwhan, et al.
Publicado: (2024) -
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
por: Nikolic, Strahinja, et al.
Publicado: (2025) -
Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
por: Choi, Hahyeon, et al.
Publicado: (2026)