On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Mingze, E, Weinan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026)
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
How Transformers Get Rich: Approximation and Dynamics Analysis
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
Task-Aware Mixture-of-Experts for Time Series Analysis
von: Wu, Xingjian, et al.
Veröffentlicht: (2025)
von: Wu, Xingjian, et al.
Veröffentlicht: (2025)
Multi-Task Reinforcement Learning with Mixture of Orthogonal Experts
von: Hendawy, Ahmed, et al.
Veröffentlicht: (2023)
von: Hendawy, Ahmed, et al.
Veröffentlicht: (2023)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
von: Zhang, Zeliang, et al.
Veröffentlicht: (2024)
On the Expressive Power of Tree-Structured Probabilistic Circuits
von: Yin, Lang, et al.
Veröffentlicht: (2024)
von: Yin, Lang, et al.
Veröffentlicht: (2024)
More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
von: Wang, Mingze, et al.
Veröffentlicht: (2026)
von: Wang, Mingze, et al.
Veröffentlicht: (2026)
Generalizable and Efficient Automated Scoring with a Knowledge-Distilled Multi-Task Mixture-of-Experts
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
Split-on-Share: Mixture of Sparse Experts for Task-Agnostic Continual Learning
von: Siddika, Fatema, et al.
Veröffentlicht: (2026)
von: Siddika, Fatema, et al.
Veröffentlicht: (2026)
FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning
von: Han, Xing, et al.
Veröffentlicht: (2026)
von: Han, Xing, et al.
Veröffentlicht: (2026)
On The Expressive Power of GNN Derivatives
von: Eitan, Yam, et al.
Veröffentlicht: (2025)
von: Eitan, Yam, et al.
Veröffentlicht: (2025)
Mixture of Experts for Recognizing Depression from Interview and Reading Tasks
von: Ilias, Loukas, et al.
Veröffentlicht: (2025)
von: Ilias, Loukas, et al.
Veröffentlicht: (2025)
Mixture of Lookup Key-Value Experts
von: Wang, Zongcheng
Veröffentlicht: (2025)
von: Wang, Zongcheng
Veröffentlicht: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
Understanding Expert Structures on Minimax Parameter Estimation in Contaminated Mixture of Experts
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
von: Yan, Fanqi, et al.
Veröffentlicht: (2024)
Weisfeiler Lehman Test on Combinatorial Complexes: Generalized Expressive Power of Topological Neural Networks
von: Chen, Jiawen, et al.
Veröffentlicht: (2026)
von: Chen, Jiawen, et al.
Veröffentlicht: (2026)
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
SPADE: Spatial Transcriptomics and Pathology Alignment Using a Mixture of Data Experts for an Expressive Latent Space
von: Redekop, Ekaterina, et al.
Veröffentlicht: (2025)
von: Redekop, Ekaterina, et al.
Veröffentlicht: (2025)
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts
von: Park, Sejik
Veröffentlicht: (2024)
von: Park, Sejik
Veröffentlicht: (2024)
MoRE: A Mixture of Low-Rank Experts for Adaptive Multi-Task Learning
von: Zhang, Dacao, et al.
Veröffentlicht: (2025)
von: Zhang, Dacao, et al.
Veröffentlicht: (2025)
MVMoE: Multi-Task Vehicle Routing Solver with Mixture-of-Experts
von: Zhou, Jianan, et al.
Veröffentlicht: (2024)
von: Zhou, Jianan, et al.
Veröffentlicht: (2024)
Mixture of Lookup Experts
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
von: Jie, Shibo, et al.
Veröffentlicht: (2025)
MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning
von: Huang, Suning, et al.
Veröffentlicht: (2024)
von: Huang, Suning, et al.
Veröffentlicht: (2024)
On the Expressive Power of Floating-Point Transformers
von: Park, Sejun, et al.
Veröffentlicht: (2026)
von: Park, Sejun, et al.
Veröffentlicht: (2026)
Expressive Power of Temporal Message Passing
von: Wałęga, Przemysław Andrzej, et al.
Veröffentlicht: (2024)
von: Wałęga, Przemysław Andrzej, et al.
Veröffentlicht: (2024)
On the Expressive Power of Contextual Relations in Transformers
von: Fraiman, Demián
Veröffentlicht: (2026)
von: Fraiman, Demián
Veröffentlicht: (2026)
Rethinking the Expressive Power of GNNs via Graph Biconnectivity
von: Zhang, Bohang, et al.
Veröffentlicht: (2023)
von: Zhang, Bohang, et al.
Veröffentlicht: (2023)
Generalizing GNNs with Tokenized Mixture of Experts
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoguang, et al.
Veröffentlicht: (2026)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
von: Wang, Jinbo, et al.
Veröffentlicht: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts
von: Wu, Shirley, et al.
Veröffentlicht: (2023)
von: Wu, Shirley, et al.
Veröffentlicht: (2023)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
von: Nguyen, Dung V., et al.
Veröffentlicht: (2025)
$μ$-Parametrization for Mixture of Experts
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
von: Małaśnicki, Jan, et al.
Veröffentlicht: (2025)
Path-Constrained Mixture-of-Experts
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
von: Gu, Zijin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
von: Wang, Mingze, et al.
Veröffentlicht: (2024) -
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
von: Liu, Yuzhi, et al.
Veröffentlicht: (2026) -
GradPower: Powering Gradients for Faster Language Model Pre-Training
von: Wang, Jinbo, et al.
Veröffentlicht: (2025) -
How Transformers Get Rich: Approximation and Dynamics Analysis
von: Wang, Mingze, et al.
Veröffentlicht: (2024) -
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
von: Jawaid, Ahad, et al.
Veröffentlicht: (2024)