RQ-MoE: Residual Quantization via Mixture of Experts for Efficient Input-Dependent Vector Compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhong, Zhengjia, Ke, Shuyan, Lin, Zaizhou, Song, Jiaqi, Lan, Hongyi, Li, Hui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025)
por: Chen, Xiaodong, et al.
Publicado: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025)
por: Li, Lujun, et al.
Publicado: (2025)
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
por: Wu, Haoyuan, et al.
Publicado: (2025)
por: Wu, Haoyuan, et al.
Publicado: (2025)
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
por: Tang, Yehui, et al.
Publicado: (2025)
por: Tang, Yehui, et al.
Publicado: (2025)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts
por: Sun, Weigao, et al.
Publicado: (2025)
por: Sun, Weigao, et al.
Publicado: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
por: Miao, Ruijie, et al.
Publicado: (2025)
por: Miao, Ruijie, et al.
Publicado: (2025)
Horseshoe Mixtures-of-Experts (HS-MoE)
por: Polson, Nick, et al.
Publicado: (2026)
por: Polson, Nick, et al.
Publicado: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
por: Wang, Wenfeng, et al.
Publicado: (2025)
por: Wang, Wenfeng, et al.
Publicado: (2025)
Hexa-MoE: Efficient and Heterogeneous-aware Training for Mixture-of-Experts
por: Luo, Shuqing, et al.
Publicado: (2024)
por: Luo, Shuqing, et al.
Publicado: (2024)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
por: Ma, Songkai, et al.
Publicado: (2025)
por: Ma, Songkai, et al.
Publicado: (2025)
RQ-GMM: Residual Quantized Gaussian Mixture Model for Multimodal Semantic Discretization in CTR Prediction
por: Tong, Ziye, et al.
Publicado: (2026)
por: Tong, Ziye, et al.
Publicado: (2026)
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
por: Takashiro, Shota, et al.
Publicado: (2026)
por: Takashiro, Shota, et al.
Publicado: (2026)
MH-MoE: Multi-Head Mixture-of-Experts
por: Huang, Shaohan, et al.
Publicado: (2024)
por: Huang, Shaohan, et al.
Publicado: (2024)
MoE-Loco: Mixture of Experts for Multitask Locomotion
por: Huang, Runhan, et al.
Publicado: (2025)
por: Huang, Runhan, et al.
Publicado: (2025)
Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
por: Hua, Yongxiang, et al.
Publicado: (2025)
por: Hua, Yongxiang, et al.
Publicado: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
por: Ludziejewski, Jan, et al.
Publicado: (2025)
por: Ludziejewski, Jan, et al.
Publicado: (2025)
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
Semi-MoE: Mixture-of-Experts meets Semi-Supervised Histopathology Segmentation
por: Vu, Nguyen Lan Vi, et al.
Publicado: (2025)
por: Vu, Nguyen Lan Vi, et al.
Publicado: (2025)
LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
por: Liu, Xinyi, et al.
Publicado: (2026)
por: Liu, Xinyi, et al.
Publicado: (2026)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
por: Neogi, Pinaki Prasad Guha, et al.
Publicado: (2025)
FT-MoE: Sustainable-learning Mixture of Experts for Fault-Tolerant Computing
por: Xiao, Wenjing, et al.
Publicado: (2025)
por: Xiao, Wenjing, et al.
Publicado: (2025)
ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
por: Ai, Mengting, et al.
Publicado: (2025)
por: Ai, Mengting, et al.
Publicado: (2025)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
por: Gu, Naibin, et al.
Publicado: (2025)
por: Gu, Naibin, et al.
Publicado: (2025)
MiLo: Efficient Quantized MoE Inference with Mixture of Low-Rank Compensators
por: Huang, Beichen, et al.
Publicado: (2025)
por: Huang, Beichen, et al.
Publicado: (2025)
Mixture of Experts (MoE): A Big Data Perspective
por: Gan, Wensheng, et al.
Publicado: (2025)
por: Gan, Wensheng, et al.
Publicado: (2025)
SDG-MoE: Signed Debate Graph Mixture-of-Experts
por: Kulibaba, Stepan, et al.
Publicado: (2026)
por: Kulibaba, Stepan, et al.
Publicado: (2026)
ECG-MoE: Mixture-of-Expert Electrocardiogram Foundation Model
por: Xu, Yuhao, et al.
Publicado: (2026)
por: Xu, Yuhao, et al.
Publicado: (2026)
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
por: Jin, In-Hwan, et al.
Publicado: (2025)
por: Jin, In-Hwan, et al.
Publicado: (2025)
GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs
por: Deng, Jianing, et al.
Publicado: (2026)
por: Deng, Jianing, et al.
Publicado: (2026)
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning
por: Do, Giang, et al.
Publicado: (2025)
por: Do, Giang, et al.
Publicado: (2025)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
por: Jin, Chao, et al.
Publicado: (2025)
por: Jin, Chao, et al.
Publicado: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
por: Qian, Yulei, et al.
Publicado: (2024)
por: Qian, Yulei, et al.
Publicado: (2024)
VA-MoE: Variables-Adaptive Mixture of Experts for Incremental Weather Forecasting
por: Chen, Hao, et al.
Publicado: (2024)
por: Chen, Hao, et al.
Publicado: (2024)
DAG-MoE: From Simple Mixture to Structural Aggregation in Mixture-of-Experts
por: Feng, Jiarui, et al.
Publicado: (2026)
por: Feng, Jiarui, et al.
Publicado: (2026)
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models
por: Aghdam, Maryam Akhavan, et al.
Publicado: (2024)
por: Aghdam, Maryam Akhavan, et al.
Publicado: (2024)
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts
por: Muzio, Alexandre, et al.
Publicado: (2024)
por: Muzio, Alexandre, et al.
Publicado: (2024)
Marco-MoE: Open Multilingual Mixture-of-Expert Language Models with Efficient Upcycling
por: Jiang, Fan, et al.
Publicado: (2026)
por: Jiang, Fan, et al.
Publicado: (2026)
Ejemplares similares
-
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
por: Chen, Xiaodong, et al.
Publicado: (2025) -
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
por: Li, Lujun, et al.
Publicado: (2025) -
Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts
por: Wu, Haoyuan, et al.
Publicado: (2025) -
Pangu Pro MoE: Mixture of Grouped Experts for Efficient Sparsity
por: Tang, Yehui, et al.
Publicado: (2025) -
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
por: Li, Yunxin, et al.
Publicado: (2024)