RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Chen, Yue, Yuxuan, Xu, Zukang, Hu, Xing, Yu, Jiangyong, Chen, Zhixuan, Zhou, Sifan, Yuan, Zhihang, Yang, Dawei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025)
by: Hu, Xing, et al.
Published: (2025)
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
by: Hu, Xing, et al.
Published: (2024)
by: Hu, Xing, et al.
Published: (2024)
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
by: Yue, Yuxuan, et al.
Published: (2025)
by: Yue, Yuxuan, et al.
Published: (2025)
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
by: Xu, Zukang, et al.
Published: (2025)
by: Xu, Zukang, et al.
Published: (2025)
MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization
by: Wang, Zhong, et al.
Published: (2026)
by: Wang, Zhong, et al.
Published: (2026)
TORQ: Two-Level Orthogonal Rotation for MXFP4 Quantization
by: Xu, Zukang, et al.
Published: (2026)
by: Xu, Zukang, et al.
Published: (2026)
MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Full Static Quantization
by: Yu, JiangYong, et al.
Published: (2025)
by: Yu, JiangYong, et al.
Published: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
by: Yu, Jiangyong, et al.
Published: (2025)
by: Yu, Jiangyong, et al.
Published: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
by: Yu, Jiangyong, et al.
Published: (2025)
by: Yu, Jiangyong, et al.
Published: (2025)
SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
by: Hu, Xing, et al.
Published: (2026)
by: Hu, Xing, et al.
Published: (2026)
Information Entropy Guided Height-aware Histogram for Quantization-friendly Pillar Feature Encoder
by: Zhou, Sifan, et al.
Published: (2024)
by: Zhou, Sifan, et al.
Published: (2024)
BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs
by: Zhao, Zhixiong, et al.
Published: (2026)
by: Zhao, Zhixiong, et al.
Published: (2026)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
by: Yue, Yuxuan, et al.
Published: (2024)
by: Yue, Yuxuan, et al.
Published: (2024)
GSQ-Tuning: Group-Shared Exponents Integer in Fully Quantized Training for LLMs On-Device Fine-tuning
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model
by: Wen, Houji, et al.
Published: (2026)
by: Wen, Houji, et al.
Published: (2026)
DLLMQuant: Quantizing Diffusion-based Large Language Models
by: Xu, Chen, et al.
Published: (2025)
by: Xu, Chen, et al.
Published: (2025)
MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
by: Hu, Zhaofeng, et al.
Published: (2024)
by: Hu, Zhaofeng, et al.
Published: (2024)
OTARo: Once Tuning for All Precisions toward Robust On-Device LLMs
by: Chen, Shaoyuan, et al.
Published: (2025)
by: Chen, Shaoyuan, et al.
Published: (2025)
NLI:Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs Inference
by: Yu, Jiangyong, et al.
Published: (2026)
by: Yu, Jiangyong, et al.
Published: (2026)
Vector Quantization Prompting for Continual Learning
by: Jiao, Li, et al.
Published: (2024)
by: Jiao, Li, et al.
Published: (2024)
R3-VAE: Reference Vector-Guided Rating Residual Quantization VAE for Generative Recommendation
by: Wan, Qiang, et al.
Published: (2026)
by: Wan, Qiang, et al.
Published: (2026)
AUV: Teaching Audio Universal Vector Quantization with Single Nested Codebook
by: Chen, Yushen, et al.
Published: (2025)
by: Chen, Yushen, et al.
Published: (2025)
A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication
by: Jiang, Xiao-Hang, et al.
Published: (2025)
by: Jiang, Xiao-Hang, et al.
Published: (2025)
Enhancing Vector Quantization with Distributional Matching: A Theoretical and Empirical Study
by: Fang, Xianghong, et al.
Published: (2025)
by: Fang, Xianghong, et al.
Published: (2025)
Channel-wise Vector Quantization
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
Does Vector Quantization Fail in Spatio-Temporal Forecasting? Exploring a Differentiable Sparse Soft-Vector Quantization Approach
by: Chen, Chao, et al.
Published: (2023)
by: Chen, Chao, et al.
Published: (2023)
Quantized Embedding Vectors for Controllable Diffusion Language Models
by: Kang, Cheng, et al.
Published: (2024)
by: Kang, Cheng, et al.
Published: (2024)
Autoregressive Speech Synthesis without Vector Quantization
by: Meng, Lingwei, et al.
Published: (2024)
by: Meng, Lingwei, et al.
Published: (2024)
Balance of Number of Embedding and their Dimensions in Vector Quantization
by: Chen, Hang, et al.
Published: (2024)
by: Chen, Hang, et al.
Published: (2024)
Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models
by: Liu, Ruikang, et al.
Published: (2025)
by: Liu, Ruikang, et al.
Published: (2025)
PillarTrack:Boosting Pillar Representation for Transformer-based 3D Single Object Tracking on Point Clouds
by: Xu, Weisheng, et al.
Published: (2024)
by: Xu, Weisheng, et al.
Published: (2024)
Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization
by: Zhou, Yuli, et al.
Published: (2026)
by: Zhou, Yuli, et al.
Published: (2026)
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
by: Duan, Bowen, et al.
Published: (2026)
by: Duan, Bowen, et al.
Published: (2026)
All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
by: Chen, Haowei, et al.
Published: (2025)
by: Chen, Haowei, et al.
Published: (2025)
Scalar Lattices and Probabilistic Shaping for Dithered Wyner-Ziv Quantization
by: Sener, Muhammed Yusuf, et al.
Published: (2025)
by: Sener, Muhammed Yusuf, et al.
Published: (2025)
Gaussian Rate-Distortion-Perception Coding and Entropy-Constrained Scalar Quantization
by: Xie, Li, et al.
Published: (2024)
by: Xie, Li, et al.
Published: (2024)
Similar Items
-
MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation Methods
by: Xu, Zukang, et al.
Published: (2025) -
MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
by: Hu, Xing, et al.
Published: (2025) -
OstQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution Fitting
by: Hu, Xing, et al.
Published: (2025) -
I-LLM: Efficient Integer-Only Inference for Fully-Quantized Low-Bit Large Language Models
by: Hu, Xing, et al.
Published: (2024) -
PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
by: Yue, Yuxuan, et al.
Published: (2025)