Accurate KV Cache Quantization with Outlier Tokens Tracing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Yi, Zhou, Yuechi, Qiu, Quantong, Li, Juntao, Xia, Qingrong, Li, Ping, Duan, Xinyu, Wang, Zhefeng, Zhang, Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915289827901440
author Su, Yi
Zhou, Yuechi
Qiu, Quantong
Li, Juntao
Xia, Qingrong
Li, Ping
Duan, Xinyu
Wang, Zhefeng
Zhang, Min
author_facet Su, Yi
Zhou, Yuechi
Qiu, Quantong
Li, Juntao
Xia, Qingrong
Li, Ping
Duan, Xinyu
Wang, Zhefeng
Zhang, Min
contents The impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10938
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accurate KV Cache Quantization with Outlier Tokens Tracing
Su, Yi
Zhou, Yuechi
Qiu, Quantong
Li, Juntao
Xia, Qingrong
Li, Ping
Duan, Xinyu
Wang, Zhefeng
Zhang, Min
Computation and Language
The impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput.
title Accurate KV Cache Quantization with Outlier Tokens Tracing
topic Computation and Language
url https://arxiv.org/abs/2505.10938