Multi-Head Mixture-of-Experts
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Xun, Huang, Shaohan, Wang, Wenhui, Wei, Furu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mixture of LoRA Experts
por: Wu, Xun, et al.
Publicado: (2024)
por: Wu, Xun, et al.
Publicado: (2024)
Textual Aesthetics in Large Language Models
por: Jiang, Lingjie, et al.
Publicado: (2024)
por: Jiang, Lingjie, et al.
Publicado: (2024)
MH-MoE: Multi-Head Mixture-of-Experts
por: Huang, Shaohan, et al.
Publicado: (2024)
por: Huang, Shaohan, et al.
Publicado: (2024)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
por: Zhang, Di, et al.
Publicado: (2025)
por: Zhang, Di, et al.
Publicado: (2025)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
por: Wang, Qibin, et al.
Publicado: (2025)
por: Wang, Qibin, et al.
Publicado: (2025)
BitNet Distillation
por: Wu, Xun, et al.
Publicado: (2025)
por: Wu, Xun, et al.
Publicado: (2025)
Text Diffusion with Reinforced Conditioning
por: Liu, Yuxuan, et al.
Publicado: (2024)
por: Liu, Yuxuan, et al.
Publicado: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
por: Lu, Xudong, et al.
Publicado: (2024)
por: Lu, Xudong, et al.
Publicado: (2024)
In-context Autoencoder for Context Compression in a Large Language Model
por: Ge, Tao, et al.
Publicado: (2023)
por: Ge, Tao, et al.
Publicado: (2023)
Routing-Free Mixture-of-Experts
por: Liu, Yilun, et al.
Publicado: (2026)
por: Liu, Yilun, et al.
Publicado: (2026)
Multilingual Routing in Mixture-of-Experts
por: Bandarkar, Lucas, et al.
Publicado: (2025)
por: Bandarkar, Lucas, et al.
Publicado: (2025)
Mixture of Heterogeneous Grouped Experts for Language Modeling
por: Ma, Zhicheng, et al.
Publicado: (2026)
por: Ma, Zhicheng, et al.
Publicado: (2026)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
por: He, Shwai, et al.
Publicado: (2025)
por: He, Shwai, et al.
Publicado: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
por: Gu, Naibin, et al.
Publicado: (2025)
por: Gu, Naibin, et al.
Publicado: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
por: Chen, Yanbei, et al.
Publicado: (2026)
por: Chen, Yanbei, et al.
Publicado: (2026)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
por: Zhuang, Haomin, et al.
Publicado: (2024)
por: Zhuang, Haomin, et al.
Publicado: (2024)
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
por: Zeng, Runjia, et al.
Publicado: (2025)
por: Zeng, Runjia, et al.
Publicado: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
por: Bershatsky, Daniel, et al.
Publicado: (2025)
por: Bershatsky, Daniel, et al.
Publicado: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
por: Hui, Tingfeng, et al.
Publicado: (2024)
por: Hui, Tingfeng, et al.
Publicado: (2024)
Black-Box On-Policy Distillation of Large Language Models
por: Ye, Tianzhu, et al.
Publicado: (2025)
por: Ye, Tianzhu, et al.
Publicado: (2025)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
por: Zhao, Guoliang, et al.
Publicado: (2025)
por: Zhao, Guoliang, et al.
Publicado: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
por: Chen, Yilong, et al.
Publicado: (2026)
por: Chen, Yilong, et al.
Publicado: (2026)
OLMoE: Open Mixture-of-Experts Language Models
por: Muennighoff, Niklas, et al.
Publicado: (2024)
por: Muennighoff, Niklas, et al.
Publicado: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
por: Krajewski, Jakub, et al.
Publicado: (2024)
por: Krajewski, Jakub, et al.
Publicado: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
por: Tejankar, Ajinkya, et al.
Publicado: (2024)
por: Tejankar, Ajinkya, et al.
Publicado: (2024)
Upcycling Large Language Models into Mixture of Experts
por: He, Ethan, et al.
Publicado: (2024)
por: He, Ethan, et al.
Publicado: (2024)
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias
por: Wang, Guorun, et al.
Publicado: (2024)
por: Wang, Guorun, et al.
Publicado: (2024)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
por: Han, Yixuan, et al.
Publicado: (2025)
por: Han, Yixuan, et al.
Publicado: (2025)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
por: Liu, Yilun, et al.
Publicado: (2025)
por: Liu, Yilun, et al.
Publicado: (2025)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
por: Kim, Gyeongman, et al.
Publicado: (2025)
por: Kim, Gyeongman, et al.
Publicado: (2025)
Mixtures of SubExperts for Large Language Continual Learning
por: Kang, Haeyong
Publicado: (2025)
por: Kang, Haeyong
Publicado: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
por: Kim, Junhyuck, et al.
Publicado: (2026)
por: Kim, Junhyuck, et al.
Publicado: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
On-Policy RL with Optimal Reward Baseline
por: Hao, Yaru, et al.
Publicado: (2025)
por: Hao, Yaru, et al.
Publicado: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Auto-ICL: In-Context Learning without Human Supervision
por: Yang, Jinghan, et al.
Publicado: (2023)
por: Yang, Jinghan, et al.
Publicado: (2023)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
por: Tang, Zhengyang, et al.
Publicado: (2024)
por: Tang, Zhengyang, et al.
Publicado: (2024)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
por: Gritsch, Nikolas, et al.
Publicado: (2024)
por: Gritsch, Nikolas, et al.
Publicado: (2024)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
por: Hallee, Logan, et al.
Publicado: (2024)
por: Hallee, Logan, et al.
Publicado: (2024)
Ejemplares similares
-
Mixture of LoRA Experts
por: Wu, Xun, et al.
Publicado: (2024) -
Textual Aesthetics in Large Language Models
por: Jiang, Lingjie, et al.
Publicado: (2024) -
MH-MoE: Multi-Head Mixture-of-Experts
por: Huang, Shaohan, et al.
Publicado: (2024) -
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
por: Zhang, Di, et al.
Publicado: (2025) -
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
por: Wang, Qibin, et al.
Publicado: (2025)