KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yuan, Aomufei, Wang, Zhiming, Miao, Ruijie, Wang, Dayu, Tian, Yuxuan, Wang, Zihan, Peng, Yebo, Wu, Yuhan, Yi, Bairen, Liu, Xin, Yang, Tong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915687995277312
author Yuan, Aomufei
Wang, Zhiming
Miao, Ruijie
Wang, Dayu
Tian, Yuxuan
Wang, Zihan
Peng, Yebo
Wu, Yuhan
Yi, Bairen
Liu, Xin
Yang, Tong
author_facet Yuan, Aomufei
Wang, Zhiming
Miao, Ruijie
Wang, Dayu
Tian, Yuxuan
Wang, Zihan
Peng, Yebo
Wu, Yuhan
Yi, Bairen
Liu, Xin
Yang, Tong
contents As the context length of current large language models (LLMs) rapidly increases, the memory demand for the Key-Value (KV) cache is becoming a bottleneck for LLM deployment and batch processing. Traditional KV cache compression methods typically involve permanently evicting or irreversibly merging "less important" tokens with low attention scores. This approach results in the unrecoverable loss of token information, which we call Contextual Amnesia, significantly degrading the model's information retrieval capability. To address this issue, we propose KVReviver, a reversible KV cache compression method based on the sketch algorithm. This method allows reconstructing compressed tokens from an additional data structure, thus enabling full-scale computation within limited memory. Experiments showed that in 2k-length contexts, it requires only 10% of KV Cache budget while maintaining identical end-to-end inference accuracy. For 32k-length contexts, it achieves equivalent or comparable accuracy ~2% accuracy loss) using merely 25% of KV Cache budget.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17917
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
Yuan, Aomufei
Wang, Zhiming
Miao, Ruijie
Wang, Dayu
Tian, Yuxuan
Wang, Zihan
Peng, Yebo
Wu, Yuhan
Yi, Bairen
Liu, Xin
Yang, Tong
Computation and Language
Artificial Intelligence
As the context length of current large language models (LLMs) rapidly increases, the memory demand for the Key-Value (KV) cache is becoming a bottleneck for LLM deployment and batch processing. Traditional KV cache compression methods typically involve permanently evicting or irreversibly merging "less important" tokens with low attention scores. This approach results in the unrecoverable loss of token information, which we call Contextual Amnesia, significantly degrading the model's information retrieval capability. To address this issue, we propose KVReviver, a reversible KV cache compression method based on the sketch algorithm. This method allows reconstructing compressed tokens from an additional data structure, thus enabling full-scale computation within limited memory. Experiments showed that in 2k-length contexts, it requires only 10% of KV Cache budget while maintaining identical end-to-end inference accuracy. For 32k-length contexts, it achieves equivalent or comparable accuracy ~2% accuracy loss) using merely 25% of KV Cache budget.
title KVReviver: Reversible KV Cache Compression with Sketch-Based Token Reconstruction
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.17917