MoE-I$^2$: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Cheng, Sui, Yang, Xiao, Jinqi, Huang, Lingyi, Gong, Yu, Duan, Yuanlin, Jia, Wenqi, Yin, Miao, Cheng, Yu, Yuan, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
MoE Pathfinder: Trajectory-driven Expert Pruning
por: Yang, Xican, et al.
Publicado: (2025)
por: Yang, Xican, et al.
Publicado: (2025)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
por: Sun, Weigao, et al.
Publicado: (2025)
por: Sun, Weigao, et al.
Publicado: (2025)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
por: Chen, Yuanteng, et al.
Publicado: (2025)
por: Chen, Yuanteng, et al.
Publicado: (2025)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025)
por: Chen, Xiaodong, et al.
Publicado: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
por: Zhang, Jihai, et al.
Publicado: (2024)
por: Zhang, Jihai, et al.
Publicado: (2024)
Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs
por: Zhou, Yixiao, et al.
Publicado: (2025)
por: Zhou, Yixiao, et al.
Publicado: (2025)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
por: Ye, Xin, et al.
Publicado: (2026)
por: Ye, Xin, et al.
Publicado: (2026)
SD-MoE: Spectral Decomposition for Effective Expert Specialization
por: Huang, Ruijun, et al.
Publicado: (2026)
por: Huang, Ruijun, et al.
Publicado: (2026)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
por: Li, Pingzhi, et al.
Publicado: (2025)
por: Li, Pingzhi, et al.
Publicado: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Mixture of Experts (MoE): A Big Data Perspective
por: Gan, Wensheng, et al.
Publicado: (2025)
por: Gan, Wensheng, et al.
Publicado: (2025)
Horseshoe Mixtures-of-Experts (HS-MoE)
por: Polson, Nick, et al.
Publicado: (2026)
por: Polson, Nick, et al.
Publicado: (2026)
$\texttt{MoE-RBench}$: Towards Building Reliable Language Models with Sparse Mixture-of-Experts
por: Chen, Guanjie, et al.
Publicado: (2024)
por: Chen, Guanjie, et al.
Publicado: (2024)
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts
por: Liao, Minwen, et al.
Publicado: (2025)
por: Liao, Minwen, et al.
Publicado: (2025)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
por: Aghdam, Maryam Akhavan, et al.
Publicado: (2024)
por: Aghdam, Maryam Akhavan, et al.
Publicado: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024)
por: Muzio, Alexandre, et al.
Publicado: (2024)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
por: Xu, Yu, et al.
Publicado: (2026)
por: Xu, Yu, et al.
Publicado: (2026)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
por: Xu, Yuhao, et al.
Publicado: (2026)
por: Xu, Yuhao, et al.
Publicado: (2026)
MoE3D: Mixture of Experts meets Multi-Modal 3D Understanding
por: Li, Yu, et al.
Publicado: (2025)
por: Li, Yu, et al.
Publicado: (2025)
ELRT: Efficient Low-Rank Training for Compact Convolutional Neural Networks
por: Sui, Yang, et al.
Publicado: (2024)
por: Sui, Yang, et al.
Publicado: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025)
por: Ma, Songkai, et al.
Publicado: (2025)
HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
por: Xiao, Jinqi, et al.
Publicado: (2023)
por: Xiao, Jinqi, et al.
Publicado: (2023)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
por: Liu, Yang, et al.
Publicado: (2026)
por: Liu, Yang, et al.
Publicado: (2026)
InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation
por: Xiao, Jinqi, et al.
Publicado: (2025)
por: Xiao, Jinqi, et al.
Publicado: (2025)
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts
por: Su, Yang, et al.
Publicado: (2025)
por: Su, Yang, et al.
Publicado: (2025)
MH-MoE: Multi-Head Mixture-of-Experts
por: Huang, Shaohan, et al.
Publicado: (2024)
por: Huang, Shaohan, et al.
Publicado: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
por: Huang, Runhan, et al.
Publicado: (2025)
por: Huang, Runhan, et al.
Publicado: (2025)
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
por: Ji, Shihao, et al.
Publicado: (2025)
por: Ji, Shihao, et al.
Publicado: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
por: Zhang, Zeliang, et al.
Publicado: (2024)
por: Zhang, Zeliang, et al.
Publicado: (2024)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
por: Zhang, Geng, et al.
Publicado: (2025)
por: Zhang, Geng, et al.
Publicado: (2025)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
por: Miao, Ruijie, et al.
Publicado: (2025)
por: Miao, Ruijie, et al.
Publicado: (2025)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
por: Wei, Tianwen, et al.
Publicado: (2024)
por: Wei, Tianwen, et al.
Publicado: (2024)
DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
por: Bai, Sikai, et al.
Publicado: (2025)
por: Bai, Sikai, et al.
Publicado: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
por: Gu, Naibin, et al.
Publicado: (2025)
por: Gu, Naibin, et al.
Publicado: (2025)
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
por: Zhang, Boyang, et al.
Publicado: (2025)
por: Zhang, Boyang, et al.
Publicado: (2025)
Ejemplares similares
-
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024) -
MoE Pathfinder: Trajectory-driven Expert Pruning
por: Yang, Xican, et al.
Publicado: (2025) -
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
por: Sun, Weigao, et al.
Publicado: (2025) -
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026) -
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
por: Chen, Yuanteng, et al.
Publicado: (2025)