STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Lee, Jaeseong, hwang, seung-won, Qiao, Aurick, Campos, Daniel F, Yao, Zhewei, He, Yuxiong |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs
par: Lee, Jaeseong, et autres
Publié: (2025)
par: Lee, Jaeseong, et autres
Publié: (2025)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
par: Lee, Jaeseong, et autres
Publié: (2025)
par: Lee, Jaeseong, et autres
Publié: (2025)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
par: Qiao, Aurick, et autres
Publié: (2024)
par: Qiao, Aurick, et autres
Publié: (2024)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
par: Cao, Mingyu, et autres
Publié: (2024)
par: Cao, Mingyu, et autres
Publié: (2024)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
par: Su, Zhaoyuan, et autres
Publié: (2026)
par: Su, Zhaoyuan, et autres
Publié: (2026)
Arctic Long Sequence Training: Scalable And Efficient Training For Multi-Million Token Sequences
par: Bekman, Stas, et autres
Publié: (2025)
par: Bekman, Stas, et autres
Publié: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
par: Koike-Akino, Toshiaki, et autres
Publié: (2025)
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training
par: Tang, Shengkun, et autres
Publié: (2026)
par: Tang, Shengkun, et autres
Publié: (2026)
MoE Pathfinder: Trajectory-driven Expert Pruning
par: Yang, Xican, et autres
Publié: (2025)
par: Yang, Xican, et autres
Publié: (2025)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
par: Zhang, Geng, et autres
Publié: (2025)
par: Zhang, Geng, et autres
Publié: (2025)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
par: Xie, Yanyue, et autres
Publié: (2024)
par: Xie, Yanyue, et autres
Publié: (2024)
Inference Scaling for Bridging Retrieval and Augmented Generation
par: Lee, Youngwon, et autres
Publié: (2024)
par: Lee, Youngwon, et autres
Publié: (2024)
CORD: Balancing COnsistency and Rank Distillation for Robust Retrieval-Augmented Generation
par: Lee, Youngwon, et autres
Publié: (2024)
par: Lee, Youngwon, et autres
Publié: (2024)
AIMER: Calibration-Free Task-Agnostic MoE Pruning
par: Liu, Zongfang, et autres
Publié: (2026)
par: Liu, Zongfang, et autres
Publié: (2026)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
par: Liao, Huanxuan, et autres
Publié: (2025)
par: Liao, Huanxuan, et autres
Publié: (2025)
Learning to Hint for Reinforcement Learning
par: Xia, Yu, et autres
Publié: (2026)
par: Xia, Yu, et autres
Publié: (2026)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
par: Kolawole, Steven, et autres
Publié: (2024)
par: Kolawole, Steven, et autres
Publié: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
par: Fu, Yao, et autres
Publié: (2025)
par: Fu, Yao, et autres
Publié: (2025)
EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
par: Liu, Zongfang, et autres
Publié: (2026)
par: Liu, Zongfang, et autres
Publié: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
par: Oliaro, Gabriele, et autres
Publié: (2024)
par: Oliaro, Gabriele, et autres
Publié: (2024)
ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning
par: Gao, Shangqian, et autres
Publié: (2025)
par: Gao, Shangqian, et autres
Publié: (2025)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
par: Wang, Ziyan, et autres
Publié: (2025)
par: Wang, Ziyan, et autres
Publié: (2025)
MoE-Sieve: Routing-Guided LoRA for Efficient MoE Fine-Tuning
par: Manzoni, Andrea
Publié: (2026)
par: Manzoni, Andrea
Publié: (2026)
Olica: Efficient Structured Pruning of Large Language Models without Retraining
par: He, Jiujun, et autres
Publié: (2025)
par: He, Jiujun, et autres
Publié: (2025)
Structured vs. Unstructured Pruning: An Exponential Gap
par: Ferre', Davide, et autres
Publié: (2026)
par: Ferre', Davide, et autres
Publié: (2026)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
par: Gu, Naibin, et autres
Publié: (2025)
par: Gu, Naibin, et autres
Publié: (2025)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
par: Li, Zichong, et autres
Publié: (2025)
par: Li, Zichong, et autres
Publié: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
par: Qin, Jiayu, et autres
Publié: (2025)
par: Qin, Jiayu, et autres
Publié: (2025)
Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
par: Xiao, Liu, et autres
Publié: (2025)
par: Xiao, Liu, et autres
Publié: (2025)
MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
par: Xia, Xinfeng, et autres
Publié: (2025)
par: Xia, Xinfeng, et autres
Publié: (2025)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
par: Muzio, Alexandre, et autres
Publié: (2024)
par: Muzio, Alexandre, et autres
Publié: (2024)
Structured Pruning for Diverse Best-of-N Reasoning Optimization
par: Nguyen, Hieu Trung, et autres
Publié: (2025)
par: Nguyen, Hieu Trung, et autres
Publié: (2025)
Deterministic Differentiable Structured Pruning for Large Language Models
par: Huang, Weiyu, et autres
Publié: (2026)
par: Huang, Weiyu, et autres
Publié: (2026)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
par: Wang, Zhaoyang, et autres
Publié: (2026)
par: Wang, Zhaoyang, et autres
Publié: (2026)
GRIN: GRadient-INformed MoE
par: Liu, Liyuan, et autres
Publié: (2024)
par: Liu, Liyuan, et autres
Publié: (2024)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
par: Goel, Raghavv, et autres
Publié: (2025)
par: Goel, Raghavv, et autres
Publié: (2025)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
par: Tang, Shengkun, et autres
Publié: (2025)
par: Tang, Shengkun, et autres
Publié: (2025)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
par: Lasby, Mike, et autres
Publié: (2025)
par: Lasby, Mike, et autres
Publié: (2025)
Demystifying When Pruning Works via Representation Hierarchies
par: He, Shwai, et autres
Publié: (2026)
par: He, Shwai, et autres
Publié: (2026)
On Pruning State-Space LLMs
par: Ghattas, Tamer, et autres
Publié: (2025)
par: Ghattas, Tamer, et autres
Publié: (2025)
Documents similaires
-
Gold-Switch: Training-Free Superposition of Slow- and Fast- Thinking LLMs
par: Lee, Jaeseong, et autres
Publié: (2025) -
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
par: Lee, Jaeseong, et autres
Publié: (2025) -
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
par: Qiao, Aurick, et autres
Publié: (2024) -
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
par: Cao, Mingyu, et autres
Publié: (2024) -
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
par: Su, Zhaoyuan, et autres
Publié: (2026)