Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Gyeongman, Chu, Gyouk, Yang, Eunho |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
par: Kim, Gyeongman, et autres
Publié: (2024)
par: Kim, Gyeongman, et autres
Publié: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
par: Kim, Junhyuck, et autres
Publié: (2026)
par: Kim, Junhyuck, et autres
Publié: (2026)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
par: Zhao, Zhongyu, et autres
Publié: (2024)
par: Zhao, Zhongyu, et autres
Publié: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)
par: Ma, Zhicheng, et autres
Publié: (2026)
Upcycling Large Language Models into Mixture of Experts
par: He, Ethan, et autres
Publié: (2024)
par: He, Ethan, et autres
Publié: (2024)
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM
par: Codefuse, et autres
Publié: (2025)
par: Codefuse, et autres
Publié: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
par: Li, Pingzhi, et autres
Publié: (2025)
par: Li, Pingzhi, et autres
Publié: (2025)
OLMoE: Open Mixture-of-Experts Language Models
par: Muennighoff, Niklas, et autres
Publié: (2024)
par: Muennighoff, Niklas, et autres
Publié: (2024)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
par: Zhao, Guoliang, et autres
Publié: (2025)
par: Zhao, Guoliang, et autres
Publié: (2025)
Multilingual Routing in Mixture-of-Experts
par: Bandarkar, Lucas, et autres
Publié: (2025)
par: Bandarkar, Lucas, et autres
Publié: (2025)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
par: Ling Team, et autres
Publié: (2025)
par: Ling Team, et autres
Publié: (2025)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
par: Herbst, Jeremy, et autres
Publié: (2026)
par: Herbst, Jeremy, et autres
Publié: (2026)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
par: Nakamura, Taishi, et autres
Publié: (2025)
par: Nakamura, Taishi, et autres
Publié: (2025)
Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
par: Han, Yixuan, et autres
Publié: (2025)
par: Han, Yixuan, et autres
Publié: (2025)
SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?
par: Zhuang, Haomin, et autres
Publié: (2024)
par: Zhuang, Haomin, et autres
Publié: (2024)
Mixtures of SubExperts for Large Language Continual Learning
par: Kang, Haeyong
Publié: (2025)
par: Kang, Haeyong
Publié: (2025)
Routing-Free Mixture-of-Experts
par: Liu, Yilun, et autres
Publié: (2026)
par: Liu, Yilun, et autres
Publié: (2026)
Multi-Head Mixture-of-Experts
par: Wu, Xun, et autres
Publié: (2024)
par: Wu, Xun, et autres
Publié: (2024)
Knowledge Localization in Mixture-of-Experts LLMs Using Cross-Lingual Inconsistency
par: Bandarkar, Lucas, et autres
Publié: (2026)
par: Bandarkar, Lucas, et autres
Publié: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
par: Pan, Bowen, et autres
Publié: (2024)
par: Pan, Bowen, et autres
Publié: (2024)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
par: Bershatsky, Daniel, et autres
Publié: (2025)
par: Bershatsky, Daniel, et autres
Publié: (2025)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
par: Zhou, Tianyi, et autres
Publié: (2026)
par: Zhou, Tianyi, et autres
Publié: (2026)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
par: Nguyen, Nam V., et autres
Publié: (2024)
par: Nguyen, Nam V., et autres
Publié: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
par: Krajewski, Jakub, et autres
Publié: (2024)
par: Krajewski, Jakub, et autres
Publié: (2024)
MoIN: Mixture of Introvert Experts to Upcycle an LLM
par: Tejankar, Ajinkya, et autres
Publié: (2024)
par: Tejankar, Ajinkya, et autres
Publié: (2024)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
par: Liu, Baihui, et autres
Publié: (2026)
par: Liu, Baihui, et autres
Publié: (2026)
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling
par: Thiombiano, Abdoul Majid O., et autres
Publié: (2025)
par: Thiombiano, Abdoul Majid O., et autres
Publié: (2025)
Graph Knowledge Distillation to Mixture of Experts
par: Rumiantsev, Pavel, et autres
Publié: (2024)
par: Rumiantsev, Pavel, et autres
Publié: (2024)
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
par: Zhou, Fan, et autres
Publié: (2024)
par: Zhou, Fan, et autres
Publié: (2024)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
par: Jiang, Juyong, et autres
Publié: (2026)
par: Jiang, Juyong, et autres
Publié: (2026)
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis
par: Li, Junzhuo, et autres
Publié: (2025)
par: Li, Junzhuo, et autres
Publié: (2025)
MobileMoE: Scaling On-Device Mixture of Experts
par: Chen, Yanbei, et autres
Publié: (2026)
par: Chen, Yanbei, et autres
Publié: (2026)
MedVAL: Toward Expert-Level Medical Text Validation with Language Models
par: Aali, Asad, et autres
Publié: (2025)
par: Aali, Asad, et autres
Publié: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
par: Pióro, Maciej, et autres
Publié: (2024)
par: Pióro, Maciej, et autres
Publié: (2024)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
par: Lou, Yuxuan, et autres
Publié: (2026)
par: Lou, Yuxuan, et autres
Publié: (2026)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
par: Xue, Fuzhao, et autres
Publié: (2024)
par: Xue, Fuzhao, et autres
Publié: (2024)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
par: Liu, Yilun, et autres
Publié: (2025)
par: Liu, Yilun, et autres
Publié: (2025)
Documents similaires
-
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
par: Kim, Gyeongman, et autres
Publié: (2024) -
Pruning and Distilling Mixture-of-Experts into Dense Language Models
par: Kim, Junhyuck, et autres
Publié: (2026) -
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
par: Zhao, Zhongyu, et autres
Publié: (2024) -
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024) -
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)