RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Chen, Yue, Yuxuan, Xu, Zukang, Hu, Xing, Yu, Jiangyong, Chen, Zhixuan, Zhou, Sifan, Yuan, Zhihang, Yang, Dawei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908352270827520
author Xu, Chen
Yue, Yuxuan
Xu, Zukang
Hu, Xing
Yu, Jiangyong
Chen, Zhixuan
Zhou, Sifan
Yuan, Zhihang
Yang, Dawei
author_facet Xu, Chen
Yue, Yuxuan
Xu, Zukang
Hu, Xing
Yu, Jiangyong
Chen, Zhixuan
Zhou, Sifan
Yuan, Zhihang
Yang, Dawei
contents RWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14x speed up.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03803
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
Xu, Chen
Yue, Yuxuan
Xu, Zukang
Hu, Xing
Yu, Jiangyong
Chen, Zhixuan
Zhou, Sifan
Yuan, Zhihang
Yang, Dawei
Machine Learning
Artificial Intelligence
RWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14x speed up.
title RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.03803