Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Rhui Dih, Wynter, Laura, Ganti, Raghu Kiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Collaboratively adding new knowledge to an LLM
von: Lee, Rhui Dih, et al.
Veröffentlicht: (2024)
von: Lee, Rhui Dih, et al.
Veröffentlicht: (2024)
Enhancing Training Efficiency Using Packing with Flash Attention
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
Efficiently Distilling LLMs for Edge Applications
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
von: Zhu, Shien, et al.
Veröffentlicht: (2025)
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
von: Li, Dengchun, et al.
Veröffentlicht: (2024)
von: Li, Dengchun, et al.
Veröffentlicht: (2024)
Upcycling Large Language Models into Mixture of Experts
von: He, Ethan, et al.
Veröffentlicht: (2024)
von: He, Ethan, et al.
Veröffentlicht: (2024)
SciDFM: A Large Language Model with Mixture-of-Experts for Science
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
von: Sun, Liangtai, et al.
Veröffentlicht: (2024)
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
von: Kim, Gyeongman, et al.
Veröffentlicht: (2025)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
von: Sukhbaatar, Sainbayar, et al.
Veröffentlicht: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jianbin, et al.
Veröffentlicht: (2025)
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
von: Liu, Boan, et al.
Veröffentlicht: (2023)
von: Liu, Boan, et al.
Veröffentlicht: (2023)
Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model
von: Li, Chong, et al.
Veröffentlicht: (2025)
von: Li, Chong, et al.
Veröffentlicht: (2025)
Mixtures of SubExperts for Large Language Continual Learning
von: Kang, Haeyong
Veröffentlicht: (2025)
von: Kang, Haeyong
Veröffentlicht: (2025)
Mix of Experts Language Model for Named Entity Recognition
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
von: Chen, Xinwei, et al.
Veröffentlicht: (2024)
Towards a Benchmark for Large Language Models for Business Process Management Tasks
von: Busch, Kiran, et al.
Veröffentlicht: (2024)
von: Busch, Kiran, et al.
Veröffentlicht: (2024)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
von: Zhao, Zhongyu, et al.
Veröffentlicht: (2024)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
von: Herbst, Jeremy, et al.
Veröffentlicht: (2026)
Mixture of Heterogeneous Grouped Experts for Language Modeling
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
von: Ma, Zhicheng, et al.
Veröffentlicht: (2026)
An Evaluation on Large Language Model Outputs: Discourse and Memorization
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2023)
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
von: Lee, Wonjun, et al.
Veröffentlicht: (2026)
If LLMs Have Human-Like Attributes, Then So Does Age of Empires II
von: de Wynter, Adrian
Veröffentlicht: (2026)
von: de Wynter, Adrian
Veröffentlicht: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2025)
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
von: Xu, Benfeng, et al.
Veröffentlicht: (2023)
Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2025)
von: Ternovtsii, Ivan, et al.
Veröffentlicht: (2025)
Large Language Models as Planning Domain Generators
von: Oswald, James, et al.
Veröffentlicht: (2024)
von: Oswald, James, et al.
Veröffentlicht: (2024)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
von: Nguyen, Nam V., et al.
Veröffentlicht: (2024)
von: Nguyen, Nam V., et al.
Veröffentlicht: (2024)
MultiPL-MoE: Multi-Programming-Lingual Extension of Large Language Models through Hybrid Mixture-of-Experts
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
RegMix: Data Mixture as Regression for Language Model Pre-training
von: Liu, Qian, et al.
Veröffentlicht: (2024)
von: Liu, Qian, et al.
Veröffentlicht: (2024)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
von: GigaChat team, et al.
Veröffentlicht: (2025)
von: GigaChat team, et al.
Veröffentlicht: (2025)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
von: Jiang, Fan, et al.
Veröffentlicht: (2026)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
Scalable Multi-Domain Adaptation of Language Models using Modular Experts
von: Schafhalter, Peter, et al.
Veröffentlicht: (2024)
von: Schafhalter, Peter, et al.
Veröffentlicht: (2024)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2025)
Is In-Context Learning Learning?
von: de Wynter, Adrian
Veröffentlicht: (2025)
von: de Wynter, Adrian
Veröffentlicht: (2025)
If Eleanor Rigby Had Met ChatGPT: A Study on Loneliness in a Post-LLM World
von: de Wynter, Adrian
Veröffentlicht: (2024)
von: de Wynter, Adrian
Veröffentlicht: (2024)
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
von: DeepSeek-AI, et al.
Veröffentlicht: (2024)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Collaboratively adding new knowledge to an LLM
von: Lee, Rhui Dih, et al.
Veröffentlicht: (2024) -
Enhancing Training Efficiency Using Packing with Flash Attention
von: Kundu, Achintya, et al.
Veröffentlicht: (2024) -
Efficiently Distilling LLMs for Edge Applications
von: Kundu, Achintya, et al.
Veröffentlicht: (2024) -
Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
von: Zhu, Shien, et al.
Veröffentlicht: (2025) -
MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
von: Li, Dengchun, et al.
Veröffentlicht: (2024)