Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information Density
Fuente:
arXiv
Guardado en:
| Autores principales: | Mi, Zhendong, Chen, Yixiao, Zhao, Pu, Yu, Xiaodong, Wang, Hao, Wang, Yanzhi, Huang, Shaoyi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Collaborative Compression for Large-Scale MoE Deployment on Edge
por: Chen, Yixiao, et al.
Publicado: (2025)
por: Chen, Yixiao, et al.
Publicado: (2025)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025)
por: Ma, Songkai, et al.
Publicado: (2025)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025)
por: Chen, Xiaodong, et al.
Publicado: (2025)
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
por: Zhu, Hengyi, et al.
Publicado: (2026)
por: Zhu, Hengyi, et al.
Publicado: (2026)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
por: Huang, Ruijun, et al.
Publicado: (2026)
por: Huang, Ruijun, et al.
Publicado: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
por: Yang, Cheng, et al.
Publicado: (2024)
por: Yang, Cheng, et al.
Publicado: (2024)
Guiding the Experts: Semantic Priors for Efficient and Focused MoE Routing
por: Min, Chengxi, et al.
Publicado: (2025)
por: Min, Chengxi, et al.
Publicado: (2025)
D$^{2}$MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
por: Wang, Haodong, et al.
Publicado: (2025)
por: Wang, Haodong, et al.
Publicado: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations
por: Jaiswal, Ajay, et al.
Publicado: (2025)
por: Jaiswal, Ajay, et al.
Publicado: (2025)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
por: Tang, Yehui, et al.
Publicado: (2025)
por: Tang, Yehui, et al.
Publicado: (2025)
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
por: Zhu, Peijun, et al.
Publicado: (2025)
por: Zhu, Peijun, et al.
Publicado: (2025)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing
por: Zhou, Hao, et al.
Publicado: (2024)
por: Zhou, Hao, et al.
Publicado: (2024)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
por: Ye, Charles, et al.
Publicado: (2026)
por: Ye, Charles, et al.
Publicado: (2026)
Harder Tasks Need More Experts: Dynamic Routing in MoE Models
por: Huang, Quzhe, et al.
Publicado: (2024)
por: Huang, Quzhe, et al.
Publicado: (2024)
Expert Routing for Communication-Efficient MoE via Finite Expert Banks
por: Salehi, Mohammad Reza Deylam, et al.
Publicado: (2026)
por: Salehi, Mohammad Reza Deylam, et al.
Publicado: (2026)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
por: Miao, Ruijie, et al.
Publicado: (2025)
por: Miao, Ruijie, et al.
Publicado: (2025)
CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering
por: Zeng, Xiyin, et al.
Publicado: (2026)
por: Zeng, Xiyin, et al.
Publicado: (2026)
BuddyMoE: Exploiting Expert Redundancy to Accelerate Memory-Constrained Mixture-of-Experts Inference
por: Wang, Yun, et al.
Publicado: (2025)
por: Wang, Yun, et al.
Publicado: (2025)
GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference
por: Han, Yu, et al.
Publicado: (2025)
por: Han, Yu, et al.
Publicado: (2025)
MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
por: Falke, Tobias, et al.
Publicado: (2026)
por: Falke, Tobias, et al.
Publicado: (2026)
MoE-nD: Per-Layer Mixture-of-Experts Routing for Multi-Axis KV Cache Compression
por: Sun, Libo, et al.
Publicado: (2026)
por: Sun, Libo, et al.
Publicado: (2026)
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
por: Wu, Juntong, et al.
Publicado: (2026)
por: Wu, Juntong, et al.
Publicado: (2026)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
por: Hu, Wentao, et al.
Publicado: (2026)
por: Hu, Wentao, et al.
Publicado: (2026)
The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise
por: Wang, Xi, et al.
Publicado: (2026)
por: Wang, Xi, et al.
Publicado: (2026)
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
por: Zhong, Shuzhang, et al.
Publicado: (2024)
por: Zhong, Shuzhang, et al.
Publicado: (2024)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
por: Shi, Long, et al.
Publicado: (2025)
por: Shi, Long, et al.
Publicado: (2025)
Delta Decompression for MoE-based LLMs Compression
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
por: Yu, Hanfei, et al.
Publicado: (2025)
por: Yu, Hanfei, et al.
Publicado: (2025)
ZipMoE: Efficient On-Device MoE Serving via Lossless Compression and Cache-Affinity Scheduling
por: Yang, Yuchen, et al.
Publicado: (2026)
por: Yang, Yuchen, et al.
Publicado: (2026)
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
por: Lv, Bo, et al.
Publicado: (2026)
por: Lv, Bo, et al.
Publicado: (2026)
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router
por: Xie, Yanyue, et al.
Publicado: (2024)
por: Xie, Yanyue, et al.
Publicado: (2024)
Layer-wise dynamic rank for compressing large language models
por: Mi, Zhendong, et al.
Publicado: (2025)
por: Mi, Zhendong, et al.
Publicado: (2025)
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
por: Wang, Liujianfu, et al.
Publicado: (2025)
por: Wang, Liujianfu, et al.
Publicado: (2025)
Expert Divergence Learning for MoE-based Language Models
por: Li, Jiaang, et al.
Publicado: (2026)
por: Li, Jiaang, et al.
Publicado: (2026)
Ejemplares similares
-
Collaborative Compression for Large-Scale MoE Deployment on Edge
por: Chen, Yixiao, et al.
Publicado: (2025) -
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025) -
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025) -
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025) -
Different Prompts, Different Ranks: Prompt-aware Dynamic Rank Selection for SVD-based LLM Compression
por: Zhu, Hengyi, et al.
Publicado: (2026)