MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jinhao, Zhang, Yunquan, Zhang, Boyang, Liu, Zeyu, Cheng, Daning |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026)
by: Ye, Xin, et al.
Published: (2026)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
by: Zhang, Jinhao Zhang Yunquan, et al.
Published: (2026)
by: Zhang, Jinhao Zhang Yunquan, et al.
Published: (2026)
Compression for Better: A General and Stable Lossless Compression Framework
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025)
by: Zhang, Jinhao, et al.
Published: (2025)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
by: Xu, Lu, et al.
Published: (2025)
by: Xu, Lu, et al.
Published: (2025)
Efficient and Robust Quantization-aware Training via Adaptive Coreset Selection
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts
by: Liu, Yuejiang, et al.
Published: (2024)
by: Liu, Yuejiang, et al.
Published: (2024)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
by: Xin, Jiayi, et al.
Published: (2025)
by: Xin, Jiayi, et al.
Published: (2025)
MoEIoU: Rethinking Bounding-Box Regression as a Mixture of Experts
by: Edula, Vinay, et al.
Published: (2026)
by: Edula, Vinay, et al.
Published: (2026)
TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion Models
by: Huang, Yushi, et al.
Published: (2023)
by: Huang, Yushi, et al.
Published: (2023)
Advancing Multimodal Large Language Models with Quantization-Aware Scale Learning for Efficient Adaptation
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
A Qualitative Test-Risk Mechanism for Scaling Behavior in Normalized Residual Networks
by: Cheng, Daning, et al.
Published: (2026)
by: Cheng, Daning, et al.
Published: (2026)
FP=xINT:Representing Neural Networks via Low-Bit Series Basis Functions
by: Zhang, Boyang, et al.
Published: (2024)
by: Zhang, Boyang, et al.
Published: (2024)
BioFact-MoE: Biologically Factorized Mixture of Experts for Vision-Language Prognostic Modeling in Hepatocellular Carcinoma
by: Yang, Junlin, et al.
Published: (2026)
by: Yang, Junlin, et al.
Published: (2026)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
by: Jia, Chenwei, et al.
Published: (2026)
by: Jia, Chenwei, et al.
Published: (2026)
Zero-Shot Quantization via Weight-Space Arithmetic
by: Solombrino, Daniele, et al.
Published: (2026)
by: Solombrino, Daniele, et al.
Published: (2026)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
by: Xia, Junhao, et al.
Published: (2025)
by: Xia, Junhao, et al.
Published: (2025)
Taming the Entropy Cliff: Variable Codebook Size Quantization for Autoregressive Visual Generation
by: Zheng, Bowen, et al.
Published: (2026)
by: Zheng, Bowen, et al.
Published: (2026)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
by: Jiang, Ruixiang, et al.
Published: (2024)
by: Jiang, Ruixiang, et al.
Published: (2024)
FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
by: Wu, Zhuguanyu, et al.
Published: (2025)
by: Wu, Zhuguanyu, et al.
Published: (2025)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
by: Teo, Rachel S. Y., et al.
Published: (2025)
by: Teo, Rachel S. Y., et al.
Published: (2025)
Patch-aware Vector Quantized Codebook Learning for Unsupervised Visual Defect Detection
by: Cheng, Qisen, et al.
Published: (2025)
by: Cheng, Qisen, et al.
Published: (2025)
Routers in Vision Mixture of Experts: An Empirical Study
by: Liu, Tianlin, et al.
Published: (2024)
by: Liu, Tianlin, et al.
Published: (2024)
ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
by: Liu, Zechun, et al.
Published: (2025)
by: Liu, Zechun, et al.
Published: (2025)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Quantization Variation: A New Perspective on Training Transformers with Low-Bit Precision
by: Huang, Xijie, et al.
Published: (2023)
by: Huang, Xijie, et al.
Published: (2023)
PA-Net: Precipitation-Adaptive Mixture-of-Experts for Long-Tail Rainfall Nowcasting
by: Xiao, Xinyu, et al.
Published: (2026)
by: Xiao, Xinyu, et al.
Published: (2026)
Exploiting Block Coordinate Descent for Cost-Effective LLM Model Training
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
Self-Supervised Quantization-Aware Knowledge Distillation
by: Zhao, Kaiqi, et al.
Published: (2024)
by: Zhao, Kaiqi, et al.
Published: (2024)
$\bf{D^3}$QE: Learning Discrete Distribution Discrepancy-aware Quantization Error for Autoregressive-Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2025)
by: Zhang, Yanran, et al.
Published: (2025)
From Sparse to Soft Mixtures of Experts
by: Puigcerver, Joan, et al.
Published: (2023)
by: Puigcerver, Joan, et al.
Published: (2023)
GC-MoE: Genomics-Guided Cell-Type-Specific Mixture of Experts for Histology-Based Single-Cell Spatial Transcriptomics
by: Shiku, Kaito, et al.
Published: (2026)
by: Shiku, Kaito, et al.
Published: (2026)
Quantized Machine Learning Models for Medical Imaging in Low-Resource Healthcare Settings
by: Kanneti, Sumanth Meenan, et al.
Published: (2026)
by: Kanneti, Sumanth Meenan, et al.
Published: (2026)
Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales
by: Pan, Shuokai, et al.
Published: (2024)
by: Pan, Shuokai, et al.
Published: (2024)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
by: Qin, Guangshuo, et al.
Published: (2026)
by: Qin, Guangshuo, et al.
Published: (2026)
Similar Items
-
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026) -
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
by: Zhang, Jinhao Zhang Yunquan, et al.
Published: (2026) -
Compression for Better: A General and Stable Lossless Compression Framework
by: Zhang, Boyang, et al.
Published: (2024) -
CALM: A CKA-Guided Adaptive Layer-Wise Modularization Framework for LLM Quantization
by: Zhang, Jinhao, et al.
Published: (2025) -
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
by: Srinivas, Sakhinana Sagar, et al.
Published: (2024)