MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hu, Xing, Chen, Zhixuan, Yang, Dawei, Xu, Zukang, Xu, Chen, Yuan, Zhihang, Zhou, Sifan, Yu, Jiangyong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
par: Zhao, Zhixiong, et autres
Publié: (2026)
par: Zhao, Zhixiong, et autres
Publié: (2026)
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
par: Xu, Chen, et autres
Publié: (2025)
par: Xu, Chen, et autres
Publié: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
par: Hu, Xing, et autres
Publié: (2025)
par: Hu, Xing, et autres
Publié: (2025)
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
par: Xu, Zukang, et autres
Publié: (2025)
par: Xu, Zukang, et autres
Publié: (2025)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
par: Xu, Zukang, et autres
Publié: (2026)
par: Xu, Zukang, et autres
Publié: (2026)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
par: Hu, Xing, et autres
Publié: (2024)
par: Hu, Xing, et autres
Publié: (2024)
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization
par: Yu, JiangYong, et autres
Publié: (2025)
par: Yu, JiangYong, et autres
Publié: (2025)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
par: Yue, Yuxuan, et autres
Publié: (2025)
par: Yue, Yuxuan, et autres
Publié: (2025)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
par: Xu, Zukang, et autres
Publié: (2025)
par: Xu, Zukang, et autres
Publié: (2025)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
par: Xu, Zukang, et autres
Publié: (2026)
par: Xu, Zukang, et autres
Publié: (2026)
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
par: Hu, Xing, et autres
Publié: (2026)
par: Hu, Xing, et autres
Publié: (2026)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
par: Huang, Yushi, et autres
Publié: (2025)
par: Huang, Yushi, et autres
Publié: (2025)
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization
par: Wang, Zhong, et autres
Publié: (2026)
par: Wang, Zhong, et autres
Publié: (2026)
Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
par: Chu, Kexin, et autres
Publié: (2025)
par: Chu, Kexin, et autres
Publié: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
par: Yu, Jiangyong, et autres
Publié: (2025)
par: Yu, Jiangyong, et autres
Publié: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
par: Yu, Jiangyong, et autres
Publié: (2025)
par: Yu, Jiangyong, et autres
Publié: (2025)
CryptoMoE: Privacy-Preserving and Scalable Mixture of Experts Inference via Balanced Expert Routing
par: Zhou, Yifan, et autres
Publié: (2025)
par: Zhou, Yifan, et autres
Publié: (2025)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
par: Zhao, Zhixiong, et autres
Publié: (2026)
par: Zhao, Zhixiong, et autres
Publié: (2026)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
par: Chen, Yuanteng, et autres
Publié: (2025)
par: Chen, Yuanteng, et autres
Publié: (2025)
MoPEQ: Mixture of Mixed Precision Quantized Experts
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
par: Chitty-Venkata, Krishna Teja, et autres
Publié: (2025)
A Replicate-and-Quantize Strategy for Plug-and-Play Load Balancing of Sparse Mixture-of-Experts LLMs
par: Liu, Zijie, et autres
Publié: (2026)
par: Liu, Zijie, et autres
Publié: (2026)
Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
par: Lu, Xudong, et autres
Publié: (2024)
par: Lu, Xudong, et autres
Publié: (2024)
QuantMoE-Bench: Examining Post-Training Quantization for Mixture-of-Experts
par: Li, Pingzhi, et autres
Publié: (2024)
par: Li, Pingzhi, et autres
Publié: (2024)
Information Entropy Guided Height-aware Histogram for Quantization-friendly Pillar Feature Encoder
par: Zhou, Sifan, et autres
Publié: (2024)
par: Zhou, Sifan, et autres
Publié: (2024)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
par: Jing, Linglin, et autres
Publié: (2025)
par: Jing, Linglin, et autres
Publié: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
par: Go, Seokjin, et autres
Publié: (2025)
par: Go, Seokjin, et autres
Publié: (2025)
$ϕ$-Balancing for Mixture-of-Experts Training
par: Chen, Lizhang, et autres
Publié: (2026)
par: Chen, Lizhang, et autres
Publié: (2026)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
par: Nguyen, Xuan-Phi, et autres
Publié: (2026)
par: Nguyen, Xuan-Phi, et autres
Publié: (2026)
MoQE: Improve Quantization Model performance via Mixture of Quantization Experts
par: Zhang, Jinhao, et autres
Publié: (2025)
par: Zhang, Jinhao, et autres
Publié: (2025)
One-for-All Does Not Work! Enhancing Vulnerability Detection by Mixture-of-Experts (MoE)
par: Yang, Xu, et autres
Publié: (2025)
par: Yang, Xu, et autres
Publié: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
par: Shen, Leyang, et autres
Publié: (2024)
par: Shen, Leyang, et autres
Publié: (2024)
Learning Explainable Stock Predictions with Tweets Using Mixture of Experts
par: Xu, Wenyan, et autres
Publié: (2025)
par: Xu, Wenyan, et autres
Publié: (2025)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
par: Yue, Yuxuan, et autres
Publié: (2024)
par: Yue, Yuxuan, et autres
Publié: (2024)
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading
par: Xu, Cheng, et autres
Publié: (2026)
par: Xu, Cheng, et autres
Publié: (2026)
MoG: Mixture of Experts for Graph-based Retrieval-Augmented Generation
par: Yuan, Zheng, et autres
Publié: (2026)
par: Yuan, Zheng, et autres
Publié: (2026)
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
par: Wang, Hongyu, et autres
Publié: (2025)
par: Wang, Hongyu, et autres
Publié: (2025)
Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
par: Yao, Jinghan, et autres
Publié: (2024)
par: Yao, Jinghan, et autres
Publié: (2024)
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
par: Dai, Damai, et autres
Publié: (2024)
par: Dai, Damai, et autres
Publié: (2024)
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
par: Jin, Peng, et autres
Publié: (2024)
par: Jin, Peng, et autres
Publié: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
par: Zhou, Sifan, et autres
Publié: (2025)
par: Zhou, Sifan, et autres
Publié: (2025)
Documents similaires
-
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
par: Zhao, Zhixiong, et autres
Publié: (2026) -
RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
par: Xu, Chen, et autres
Publié: (2025) -
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
par: Hu, Xing, et autres
Publié: (2025) -
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
par: Xu, Zukang, et autres
Publié: (2025) -
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
par: Xu, Zukang, et autres
Publié: (2026)