More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiebin, Zhu, Dawei, Song, Yifan, Wu, Wenhao, Kuang, Chuqiao, Li, Xiaoguang, Shang, Lifeng, Liu, Qun, Li, Sujian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912238874394624
author Zhang, Jiebin
Zhu, Dawei
Song, Yifan
Wu, Wenhao
Kuang, Chuqiao
Li, Xiaoguang
Shang, Lifeng
Liu, Qun
Li, Sujian
author_facet Zhang, Jiebin
Zhu, Dawei
Song, Yifan
Wu, Wenhao
Kuang, Chuqiao
Li, Xiaoguang
Shang, Lifeng
Liu, Qun
Li, Sujian
contents As large language models (LLMs) process increasing context windows, the memory usage of KV cache has become a critical bottleneck during inference. The mainstream KV compression methods, including KV pruning and KV quantization, primarily focus on either token or precision dimension separately. However, these works leaving the trade-off between these two orthogonal dimensions largely under-explored. In this paper, we comprehensively investigate the token-precision trade-off in KV cache compression.Experiments demonstrate that storing more tokens in the KV cache with lower precision,a strategy we term quantized pruning, can significantly enhance the long-context performance of LLMs. In-depth analysis of the token-precision trade-off across key aspects demonstrates that, quantized pruning achieves substantial improvements in retrieval-related tasks and consistently performs well across varying input lengths. Furthermore, quantized pruning demonstrates notable stability and effectiveness across different KV pruning methods, quantization strategies, and model scales. These findings offer valuable insights into optimizing KV cache compression through balanced token-precision trade-off strategies. Our code is available at https://github.com/zhzihao/QPruningKV.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
Zhang, Jiebin
Zhu, Dawei
Song, Yifan
Wu, Wenhao
Kuang, Chuqiao
Li, Xiaoguang
Shang, Lifeng
Liu, Qun
Li, Sujian
Computation and Language
As large language models (LLMs) process increasing context windows, the memory usage of KV cache has become a critical bottleneck during inference. The mainstream KV compression methods, including KV pruning and KV quantization, primarily focus on either token or precision dimension separately. However, these works leaving the trade-off between these two orthogonal dimensions largely under-explored. In this paper, we comprehensively investigate the token-precision trade-off in KV cache compression.Experiments demonstrate that storing more tokens in the KV cache with lower precision,a strategy we term quantized pruning, can significantly enhance the long-context performance of LLMs. In-depth analysis of the token-precision trade-off across key aspects demonstrates that, quantized pruning achieves substantial improvements in retrieval-related tasks and consistently performs well across varying input lengths. Furthermore, quantized pruning demonstrates notable stability and effectiveness across different KV pruning methods, quantization strategies, and model scales. These findings offer valuable insights into optimizing KV cache compression through balanced token-precision trade-off strategies. Our code is available at https://github.com/zhzihao/QPruningKV.
title More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression
topic Computation and Language
url https://arxiv.org/abs/2412.12706