Pruning and Distilling Mixture-of-Experts into Dense Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Junhyuck, Yun, Jihun, Kim, Haechan, Kim, Gyeongman, Bae, Joonghyun, Cho, Jaewoong |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
par: Kim, Gyeongman, et autres
Publié: (2025)
par: Kim, Gyeongman, et autres
Publié: (2025)
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
par: Kim, Gyeongman, et autres
Publié: (2024)
par: Kim, Gyeongman, et autres
Publié: (2024)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
par: Kim, Jaechang, et autres
Publié: (2024)
par: Kim, Jaechang, et autres
Publié: (2024)
Raon-Speech Technical Report
par: Kim, Beomsoo, et autres
Publié: (2026)
par: Kim, Beomsoo, et autres
Publié: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning
par: Kim, Minkyu, et autres
Publié: (2026)
par: Kim, Minkyu, et autres
Publié: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
par: Pan, Bowen, et autres
Publié: (2024)
par: Pan, Bowen, et autres
Publié: (2024)
Surrogate modeling for interpreting black-box LLMs in medical predictions
par: Han, Changho, et autres
Publié: (2026)
par: Han, Changho, et autres
Publié: (2026)
DistiLLM: Towards Streamlined Distillation for Large Language Models
par: Ko, Jongwoo, et autres
Publié: (2024)
par: Ko, Jongwoo, et autres
Publié: (2024)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
par: Park, Dongmin, et autres
Publié: (2024)
par: Park, Dongmin, et autres
Publié: (2024)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models
par: Jiang, Juyong, et autres
Publié: (2026)
par: Jiang, Juyong, et autres
Publié: (2026)
Compact Language Models via Pruning and Knowledge Distillation
par: Muralidharan, Saurav, et autres
Publié: (2024)
par: Muralidharan, Saurav, et autres
Publié: (2024)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
par: Ho, Namgyu, et autres
Publié: (2024)
par: Ho, Namgyu, et autres
Publié: (2024)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
par: Kim, Taehyun, et autres
Publié: (2024)
par: Kim, Taehyun, et autres
Publié: (2024)
Mixture of Heterogeneous Grouped Experts for Language Modeling
par: Ma, Zhicheng, et autres
Publié: (2026)
par: Ma, Zhicheng, et autres
Publié: (2026)
Upcycling Large Language Models into Mixture of Experts
par: He, Ethan, et autres
Publié: (2024)
par: He, Ethan, et autres
Publié: (2024)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
par: Lee, Keon, et autres
Publié: (2024)
par: Lee, Keon, et autres
Publié: (2024)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
par: Hui, Tingfeng, et autres
Publié: (2024)
par: Hui, Tingfeng, et autres
Publié: (2024)
OLMoE: Open Mixture-of-Experts Language Models
par: Muennighoff, Niklas, et autres
Publié: (2024)
par: Muennighoff, Niklas, et autres
Publié: (2024)
Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
par: Song, Yeongwoo, et autres
Publié: (2025)
par: Song, Yeongwoo, et autres
Publié: (2025)
SDMPrune: Self-Distillation MLP Pruning for Efficient Large Language Models
par: Zhu, Hourun, et autres
Publié: (2025)
par: Zhu, Hourun, et autres
Publié: (2025)
Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
par: Nakamura, Taishi, et autres
Publié: (2025)
par: Nakamura, Taishi, et autres
Publié: (2025)
Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
par: Choi, Minsik, et autres
Publié: (2025)
par: Choi, Minsik, et autres
Publié: (2025)
Non-linear Interventions on Large Language Models
par: Kim, Sangwoo
Publié: (2026)
par: Kim, Sangwoo
Publié: (2026)
Self-Training Elicits Concise Reasoning in Large Language Models
par: Munkhbat, Tergel, et autres
Publié: (2025)
par: Munkhbat, Tergel, et autres
Publié: (2025)
Nevermind: Instruction Override and Moderation in Large Language Models
par: Kim, Edward
Publié: (2024)
par: Kim, Edward
Publié: (2024)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
par: Jung, Hee-Jun, et autres
Publié: (2022)
par: Jung, Hee-Jun, et autres
Publié: (2022)
Mixtures of SubExperts for Large Language Continual Learning
par: Kang, Haeyong
Publié: (2025)
par: Kang, Haeyong
Publié: (2025)
FactorLLM: Factorizing Knowledge via Mixture of Experts for Large Language Models
par: Zhao, Zhongyu, et autres
Publié: (2024)
par: Zhao, Zhongyu, et autres
Publié: (2024)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
par: Lee, Gihun, et autres
Publié: (2024)
par: Lee, Gihun, et autres
Publié: (2024)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
par: Cho, Junmo, et autres
Publié: (2026)
par: Cho, Junmo, et autres
Publié: (2026)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
par: Nguyen, Nam V., et autres
Publié: (2024)
par: Nguyen, Nam V., et autres
Publié: (2024)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
par: Prahlad, Deeksha, et autres
Publié: (2025)
par: Prahlad, Deeksha, et autres
Publié: (2025)
Probing Semantic Routing in Large Mixture-of-Expert Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
Routing-Free Mixture-of-Experts
par: Liu, Yilun, et autres
Publié: (2026)
par: Liu, Yilun, et autres
Publié: (2026)
Multi-Head Mixture-of-Experts
par: Wu, Xun, et autres
Publié: (2024)
par: Wu, Xun, et autres
Publié: (2024)
Multilingual Routing in Mixture-of-Experts
par: Bandarkar, Lucas, et autres
Publié: (2025)
par: Bandarkar, Lucas, et autres
Publié: (2025)
Rethinking the Role of Proxy Rewards in Language Model Alignment
par: Kim, Sungdong, et autres
Publié: (2024)
par: Kim, Sungdong, et autres
Publié: (2024)
Documents similaires
-
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models
par: Kim, Gyeongman, et autres
Publié: (2025) -
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning
par: Kim, Gyeongman, et autres
Publié: (2024) -
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
par: Kim, Jaechang, et autres
Publié: (2024) -
Raon-Speech Technical Report
par: Kim, Beomsoo, et autres
Publié: (2026) -
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)