SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Shengkun, Wang, Zekun, Zheng, Bo, Wang, Liangyu, Men, Rui, Zhang, Siqi, Yuan, Xiulong, Qiu, Zihan, Shen, Zhiqiang, Liu, Dayiheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)
by: Wang, Liangyu, et al.
Published: (2026)
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
AIMER: Calibration-Free Task-Agnostic MoE Pruning
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
by: Guo, Bowei, et al.
Published: (2025)
by: Guo, Bowei, et al.
Published: (2025)
Sink-Aware Pruning for Diffusion Language Models
by: Myrzakhan, Aidar, et al.
Published: (2026)
by: Myrzakhan, Aidar, et al.
Published: (2026)
GW-MoE: Resolving Uncertainty in MoE Router with Global Workspace Theory
by: Wu, Haoze, et al.
Published: (2024)
by: Wu, Haoze, et al.
Published: (2024)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
by: Wang, Bo, et al.
Published: (2026)
by: Wang, Bo, et al.
Published: (2026)
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
Diff-ES: Stage-wise Structural Diffusion Pruning via Evolutionary Search
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
SlimLLM: Accurate Structured Pruning for Large Language Models
by: Guo, Jialong, et al.
Published: (2025)
by: Guo, Jialong, et al.
Published: (2025)
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
by: Lu, Xin, et al.
Published: (2024)
by: Lu, Xin, et al.
Published: (2024)
DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-Experts
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Making Pre-trained Language Models Better Continual Few-Shot Relation Extractors
by: Ma, Shengkun, et al.
Published: (2024)
by: Ma, Shengkun, et al.
Published: (2024)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Qwen3 Technical Report
by: Yang, An, et al.
Published: (2025)
by: Yang, An, et al.
Published: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
by: Qiu, Zihan, et al.
Published: (2025)
by: Qiu, Zihan, et al.
Published: (2025)
Qwen2.5-Coder Technical Report
by: Hui, Binyuan, et al.
Published: (2024)
by: Hui, Binyuan, et al.
Published: (2024)
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
MoE Pathfinder: Trajectory-driven Expert Pruning
by: Yang, Xican, et al.
Published: (2025)
by: Yang, Xican, et al.
Published: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
by: Wang, Peng, et al.
Published: (2024)
by: Wang, Peng, et al.
Published: (2024)
Qwen2.5 Technical Report
by: Qwen, et al.
Published: (2024)
by: Qwen, et al.
Published: (2024)
Qwen3-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
SlimGPT: Layer-wise Structured Pruning for Large Language Models
by: Ling, Gui, et al.
Published: (2024)
by: Ling, Gui, et al.
Published: (2024)
ECo-MoE: Embodiment-Conditioned Mixture of Experts Increases the Evolvability of Robots
by: Wang, Yibin, et al.
Published: (2026)
by: Wang, Yibin, et al.
Published: (2026)
Exploring 3D Dataset Pruning
by: Zhao, Xiaohan, et al.
Published: (2026)
by: Zhao, Xiaohan, et al.
Published: (2026)
Qwen2.5-1M Technical Report
by: Yang, An, et al.
Published: (2025)
by: Yang, An, et al.
Published: (2025)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
by: Zhang, Geng, et al.
Published: (2025)
by: Zhang, Geng, et al.
Published: (2025)
LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
by: Zhu, Tong, et al.
Published: (2024)
by: Zhu, Tong, et al.
Published: (2024)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
by: Zhang, Jiyuan, et al.
Published: (2026)
by: Zhang, Jiyuan, et al.
Published: (2026)
Bi-Mamba: Towards Accurate 1-Bit State Space Models
by: Tang, Shengkun, et al.
Published: (2024)
by: Tang, Shengkun, et al.
Published: (2024)
BiGain: Unified Token Compression for Joint Generation and Classification
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
by: Miao, Ruijie, et al.
Published: (2025)
by: Miao, Ruijie, et al.
Published: (2025)
Similar Items
-
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026) -
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
by: Liu, Zongfang, et al.
Published: (2026) -
AIMER: Calibration-Free Task-Agnostic MoE Pruning
by: Liu, Zongfang, et al.
Published: (2026) -
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
by: Li, Zichong, et al.
Published: (2025) -
MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining Dynamics
by: Guo, Bowei, et al.
Published: (2025)