Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Enshu, Zhu, Junyi, Lin, Zinan, Ning, Xuefei, Blaschko, Matthew B., Yan, Shengen, Dai, Guohao, Yang, Huazhong, Wang, Yu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
por: Liu, Enshu, et al.
Publicado: (2024)
por: Liu, Enshu, et al.
Publicado: (2024)
Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
por: Liu, Enshu, et al.
Publicado: (2025)
por: Liu, Enshu, et al.
Publicado: (2025)
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
por: Zhao, Tianchen, et al.
Publicado: (2024)
por: Zhao, Tianchen, et al.
Publicado: (2024)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
por: Zhao, Tianchen, et al.
Publicado: (2024)
por: Zhao, Tianchen, et al.
Publicado: (2024)
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
por: Lin, Zinan, et al.
Publicado: (2025)
por: Lin, Zinan, et al.
Publicado: (2025)
Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching
por: Liu, Enshu, et al.
Publicado: (2024)
por: Liu, Enshu, et al.
Publicado: (2024)
NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization
por: Liu, Enshu, et al.
Publicado: (2026)
por: Liu, Enshu, et al.
Publicado: (2026)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
por: Fu, Tianyu, et al.
Publicado: (2025)
por: Fu, Tianyu, et al.
Publicado: (2025)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
por: Ning, Xuefei, et al.
Publicado: (2024)
por: Ning, Xuefei, et al.
Publicado: (2024)
FlashEval: Towards Fast and Accurate Evaluation of Text-to-image Diffusion Generative Models
por: Zhao, Lin, et al.
Publicado: (2024)
por: Zhao, Lin, et al.
Publicado: (2024)
Evaluating Quantized Large Language Models
por: Li, Shiyao, et al.
Publicado: (2024)
por: Li, Shiyao, et al.
Publicado: (2024)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
por: Lu, Xudong, et al.
Publicado: (2024)
por: Lu, Xudong, et al.
Publicado: (2024)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
por: Wang, Luning, et al.
Publicado: (2024)
por: Wang, Luning, et al.
Publicado: (2024)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
por: Ning, Xuefei, et al.
Publicado: (2023)
por: Ning, Xuefei, et al.
Publicado: (2023)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
por: Zhou, Yixiao, et al.
Publicado: (2025)
por: Zhou, Yixiao, et al.
Publicado: (2025)
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2024)
por: Chowdhury, Mohammed Nowaz Rabbani, et al.
Publicado: (2024)
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
por: Guo, Hongcheng, et al.
Publicado: (2025)
por: Guo, Hongcheng, et al.
Publicado: (2025)
Mixture-of-Linear-Experts for Long-term Time Series Forecasting
por: Ni, Ronghao, et al.
Publicado: (2023)
por: Ni, Ronghao, et al.
Publicado: (2023)
SparseDoctor: Towards Efficient Chat Doctor with Mixture of Experts Enhanced Large Language Models
por: Zhang, Jianbin, et al.
Publicado: (2025)
por: Zhang, Jianbin, et al.
Publicado: (2025)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
por: Kim, Junhyuck, et al.
Publicado: (2026)
por: Kim, Junhyuck, et al.
Publicado: (2026)
MBQ: Modality-Balanced Quantization for Large Vision-Language Models
por: Li, Shiyao, et al.
Publicado: (2024)
por: Li, Shiyao, et al.
Publicado: (2024)
Accelerating Mixture-of-Expert Inference with Adaptive Expert Split Mechanism
por: Yan, Jiaming, et al.
Publicado: (2025)
por: Yan, Jiaming, et al.
Publicado: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
por: Jiang, Yinsicheng, et al.
Publicado: (2025)
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
por: Jiang, Yinsicheng, et al.
Publicado: (2024)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
por: Benazir, Afsara, et al.
Publicado: (2026)
por: Benazir, Afsara, et al.
Publicado: (2026)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
por: Pan, Bowen, et al.
Publicado: (2024)
por: Pan, Bowen, et al.
Publicado: (2024)
Speculating Experts Accelerates Inference for Mixture-of-Experts
por: Madan, Vivan, et al.
Publicado: (2026)
por: Madan, Vivan, et al.
Publicado: (2026)
A Survey on Efficient Inference for Large Language Models
por: Zhou, Zixuan, et al.
Publicado: (2024)
por: Zhou, Zixuan, et al.
Publicado: (2024)
ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
por: He, Xin, et al.
Publicado: (2024)
por: He, Xin, et al.
Publicado: (2024)
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
por: Nguyen, Dung V., et al.
Publicado: (2025)
por: Nguyen, Dung V., et al.
Publicado: (2025)
Jaccard Metric Losses: Optimizing the Jaccard Index with Soft Labels
por: Wang, Zifu, et al.
Publicado: (2023)
por: Wang, Zifu, et al.
Publicado: (2023)
Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony
por: Wang, Shaoyu, et al.
Publicado: (2025)
por: Wang, Shaoyu, et al.
Publicado: (2025)
Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts Models
por: Hu, Wentao, et al.
Publicado: (2025)
por: Hu, Wentao, et al.
Publicado: (2025)
CRAFT: Fine-Grained Cost-Aware Expert Replication For Efficient Mixture-of-Experts Serving
por: Zhao, Adrian, et al.
Publicado: (2026)
por: Zhao, Adrian, et al.
Publicado: (2026)
MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
por: Zhao, Yushu, et al.
Publicado: (2025)
por: Zhao, Yushu, et al.
Publicado: (2025)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
por: Chu, Kexin, et al.
Publicado: (2025)
por: Chu, Kexin, et al.
Publicado: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
por: Lin, Haoran, et al.
Publicado: (2025)
por: Lin, Haoran, et al.
Publicado: (2025)
Ejemplares similares
-
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better
por: Liu, Enshu, et al.
Publicado: (2024) -
Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score Distillation
por: Liu, Enshu, et al.
Publicado: (2025) -
MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization
por: Zhao, Tianchen, et al.
Publicado: (2024) -
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
por: Zhao, Tianchen, et al.
Publicado: (2024) -
Latent Zoning Network: A Unified Principle for Generative Modeling, Representation Learning, and Classification
por: Lin, Zinan, et al.
Publicado: (2025)