Harder Tasks Need More Experts: Dynamic Routing in MoE Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Quzhe, An, Zhenwei, Zhuang, Nan, Tao, Mingxu, Zhang, Chen, Jin, Yang, Xu, Kun, Chen, Liwei, Huang, Songfang, Feng, Yansong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Probing Multimodal Large Language Models for Global and Local Semantic Representations
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
Unlocking the Potential of Model Merging for Low-Resource Languages
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
von: Tao, Mingxu, et al.
Veröffentlicht: (2024)
Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
MC$^2$: Towards Transparent and Culturally-Aware NLP for Minority Languages in China
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
von: Zhang, Chen, et al.
Veröffentlicht: (2023)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)
Only One Relation Possible? Modeling the Ambiguity in Event Temporal Relation Extraction
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
von: Xu, Yuqi, et al.
Veröffentlicht: (2026)
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
von: Hua, Yongxiang, et al.
Veröffentlicht: (2025)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
von: Wu, Juntong, et al.
Veröffentlicht: (2026)
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
von: Ma, Haiyue, et al.
Veröffentlicht: (2025)
Advancing Expert Specialization for Better MoE
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
von: Guo, Hongcan, et al.
Veröffentlicht: (2025)
SiftMoE: Similarity-Aware Energy-Efficient Expert Selection for Wireless Distributed MoE Inference
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
MoE Lens -- An Expert Is All You Need
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2026)
von: Chaudhari, Marmik, et al.
Veröffentlicht: (2026)
What Kinds of Tokens Benefit from Distant Text? An Analysis on Long Context Language Modeling
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
von: Hu, Yutong, et al.
Veröffentlicht: (2024)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
von: Luo, Kangcheng, et al.
Veröffentlicht: (2025)
von: Luo, Kangcheng, et al.
Veröffentlicht: (2025)
JUREX-4E: Juridical Expert-Annotated Four-Element Knowledge Base for Legal Reasoning
von: Liu, Huanghai, et al.
Veröffentlicht: (2025)
von: Liu, Huanghai, et al.
Veröffentlicht: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
von: Huang, Yuegui, et al.
Veröffentlicht: (2026)
von: Huang, Yuegui, et al.
Veröffentlicht: (2026)
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
von: Salehi, Mohammad Reza Deylam, et al.
Veröffentlicht: (2026)
von: Salehi, Mohammad Reza Deylam, et al.
Veröffentlicht: (2026)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
von: Aghdam, Maryam Akhavan, et al.
Veröffentlicht: (2024)
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2023)
von: Jin, Yang, et al.
Veröffentlicht: (2023)
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
von: Hu, Rizhen, et al.
Veröffentlicht: (2026)
von: Hu, Rizhen, et al.
Veröffentlicht: (2026)
CARL-MoE: Communication-Aware Adaptive Routing with Load-Balanced Expert Parallelism for Efficient Mixture-of-Experts Training
von: Jin, Haopeng
Veröffentlicht: (2026)
von: Jin, Haopeng
Veröffentlicht: (2026)
Pyramidal Flow Matching for Efficient Video Generative Modeling
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
von: Jin, In-Hwan, et al.
Veröffentlicht: (2025)
von: Jin, In-Hwan, et al.
Veröffentlicht: (2025)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
von: Xu, Yuhao, et al.
Veröffentlicht: (2026)
Orchestrating Heterogeneous Experts: A Scalable MoE Framework with Anisotropy-Preserving Fusion
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
von: Falke, Tobias, et al.
Veröffentlicht: (2026)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
von: Min, Chengxi, et al.
Veröffentlicht: (2025)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
MH-MoE: Multi-Head Mixture-of-Experts
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
von: Huang, Shaohan, et al.
Veröffentlicht: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
von: Huang, Runhan, et al.
Veröffentlicht: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information Density
von: Mi, Zhendong, et al.
Veröffentlicht: (2026)
von: Mi, Zhendong, et al.
Veröffentlicht: (2026)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
von: Wang, Liujianfu, et al.
Veröffentlicht: (2025)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
von: Lv, Bo, et al.
Veröffentlicht: (2026)
von: Lv, Bo, et al.
Veröffentlicht: (2026)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
von: Huang, Ruijun, et al.
Veröffentlicht: (2026)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
von: Han, Yu, et al.
Veröffentlicht: (2025)
von: Han, Yu, et al.
Veröffentlicht: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Probing Multimodal Large Language Models for Global and Local Semantic Representations
von: Tao, Mingxu, et al.
Veröffentlicht: (2024) -
Unlocking the Potential of Model Merging for Low-Resource Languages
von: Tao, Mingxu, et al.
Veröffentlicht: (2024) -
Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
von: Hu, Yutong, et al.
Veröffentlicht: (2024) -
MC$^2$: Towards Transparent and Culturally-Aware NLP for Minority Languages in China
von: Zhang, Chen, et al.
Veröffentlicht: (2023) -
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
von: Wu, Haoyuan, et al.
Veröffentlicht: (2025)