Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhao, Jinze, Wang, Peihao, Wang, Zhangyang |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
par: Zhao, Jinze, et autres
Publié: (2024)
par: Zhao, Jinze, et autres
Publié: (2024)
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
par: Yang, Junjie, et autres
Publié: (2023)
par: Yang, Junjie, et autres
Publié: (2023)
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
par: Wang, Peihao, et autres
Publié: (2025)
par: Wang, Peihao, et autres
Publié: (2025)
Position: Weight Space Should Be a First-Class Generative AI Modality
par: Wang, Zhangyang, et autres
Publié: (2026)
par: Wang, Zhangyang, et autres
Publié: (2026)
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
par: Cai, Ruisi, et autres
Publié: (2024)
par: Cai, Ruisi, et autres
Publié: (2024)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2024)
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2024)
Polynomial Width is Sufficient for Set Representation with High-dimensional Features
par: Wang, Peihao, et autres
Publié: (2023)
par: Wang, Peihao, et autres
Publié: (2023)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
par: Nguyen, Dung V., et autres
Publié: (2025)
par: Nguyen, Dung V., et autres
Publié: (2025)
Neural Networks with Sparse Activation Induced by Large Bias: Tighter Analysis with Bias-Generalized NTK
par: Yang, Hongru, et autres
Publié: (2023)
par: Yang, Hongru, et autres
Publié: (2023)
On the Role of Discrete Representation in Sparse Mixture of Experts
par: Do, Giang, et autres
Publié: (2024)
par: Do, Giang, et autres
Publié: (2024)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
par: Zhao, Yushu, et autres
Publié: (2025)
par: Zhao, Yushu, et autres
Publié: (2025)
Selective Sinkhorn Routing for Improved Sparse Mixture of Experts
par: Nguyen, Duc Anh, et autres
Publié: (2025)
par: Nguyen, Duc Anh, et autres
Publié: (2025)
Improving Routing in Sparse Mixture of Experts with Graph of Tokens
par: Nguyen, Tam, et autres
Publié: (2025)
par: Nguyen, Tam, et autres
Publié: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
par: Chen, Shengzhuang, et autres
Publié: (2025)
par: Chen, Shengzhuang, et autres
Publié: (2025)
Generalizing GNNs with Tokenized Mixture of Experts
par: Guo, Xiaoguang, et autres
Publié: (2026)
par: Guo, Xiaoguang, et autres
Publié: (2026)
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
par: Liu, Enshu, et autres
Publié: (2024)
par: Liu, Enshu, et autres
Publié: (2024)
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
par: Pan, Xinglin, et autres
Publié: (2025)
par: Pan, Xinglin, et autres
Publié: (2025)
When Do Graph Foundation Models Transfer? A Data-Centric Theory
par: Zhu, Jiajun, et autres
Publié: (2026)
par: Zhu, Jiajun, et autres
Publié: (2026)
Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing
par: Wang, Peihao, et autres
Publié: (2024)
par: Wang, Peihao, et autres
Publié: (2024)
$\nabla$-Reasoner: LLM Reasoning via Test-Time Gradient Descent in Latent Space
par: Wang, Peihao, et autres
Publié: (2026)
par: Wang, Peihao, et autres
Publié: (2026)
Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
par: Jaiswal, Ajay, et autres
Publié: (2025)
par: Jaiswal, Ajay, et autres
Publié: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
par: Zhang, Zeliang, et autres
Publié: (2024)
par: Zhang, Zeliang, et autres
Publié: (2024)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
par: Nguyen-Nhat, Minh-Khoi, et autres
Publié: (2025)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
par: Nikolic, Strahinja, et autres
Publié: (2025)
par: Nikolic, Strahinja, et autres
Publié: (2025)
Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts
par: Ahrac, Sagi, et autres
Publié: (2026)
par: Ahrac, Sagi, et autres
Publié: (2026)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
par: Park, Sejik
Publié: (2024)
par: Park, Sejik
Publié: (2024)
Mixture of Lookup Key-Value Experts
par: Wang, Zongcheng
Publié: (2025)
par: Wang, Zongcheng
Publié: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
par: Panda, Ashwinee, et autres
Publié: (2025)
par: Panda, Ashwinee, et autres
Publié: (2025)
CoSMoEs: Compact Sparse Mixture of Experts
par: Huber, Patrick, et autres
Publié: (2025)
par: Huber, Patrick, et autres
Publié: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
par: Muzio, Alexandre, et autres
Publié: (2024)
par: Muzio, Alexandre, et autres
Publié: (2024)
From Sparse to Soft Mixtures of Experts
par: Puigcerver, Joan, et autres
Publié: (2023)
par: Puigcerver, Joan, et autres
Publié: (2023)
Split-on-Share: Mixture of Sparse Experts for Task-Agnostic Continual Learning
par: Siddika, Fatema, et autres
Publié: (2026)
par: Siddika, Fatema, et autres
Publié: (2026)
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts
par: Nguyen, Huy, et autres
Publié: (2023)
par: Nguyen, Huy, et autres
Publié: (2023)
DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices
par: Song, Chenyang, et autres
Publié: (2026)
par: Song, Chenyang, et autres
Publié: (2026)
When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Models
par: Yoon, Youngsik, et autres
Publié: (2026)
par: Yoon, Youngsik, et autres
Publié: (2026)
Soft-to-Hard Routing in Sparse Mixture-of-Experts Models
par: Rastegar, Reza
Publié: (2026)
par: Rastegar, Reza
Publié: (2026)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
par: Gao, Yuting, et autres
Publié: (2025)
par: Gao, Yuting, et autres
Publié: (2025)
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
par: Jiang, Yukun, et autres
Publié: (2026)
par: Jiang, Yukun, et autres
Publié: (2026)
Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2026)
par: Chowdhury, Mohammed Nowaz Rabbani, et autres
Publié: (2026)
Documents similaires
-
Sparse Mixture-of-Experts for Compositional Generalization: Empirical Evidence and Theoretical Foundations of Optimal Sparsity
par: Zhao, Jinze, et autres
Publié: (2024) -
Meta ControlNet: Enhancing Task Adaptation via Meta Learning
par: Yang, Junjie, et autres
Publié: (2023) -
Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning
par: Wang, Peihao, et autres
Publié: (2025) -
Position: Weight Space Should Be a First-Class Generative AI Modality
par: Wang, Zhangyang, et autres
Publié: (2026) -
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
par: Cai, Ruisi, et autres
Publié: (2024)