R-KV: Redundancy-aware KV Cache Compression for Reasoning Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Zefan, Xiao, Wen, Sun, Hanshi, Luo, Cheng, Zhang, Yikai, Wan, Ke, Li, Yucheng, Zhou, Yeyang, Chang, Li-Wen, Gu, Jiuxiang, Dong, Zhen, Anandkumar, Anima, Asi, Abedelkadir, Hu, Junjie
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911390815485952
author Cai, Zefan
Xiao, Wen
Sun, Hanshi
Luo, Cheng
Zhang, Yikai
Wan, Ke
Li, Yucheng
Zhou, Yeyang
Chang, Li-Wen
Gu, Jiuxiang
Dong, Zhen
Anandkumar, Anima
Asi, Abedelkadir
Hu, Junjie
author_facet Cai, Zefan
Xiao, Wen
Sun, Hanshi
Luo, Cheng
Zhang, Yikai
Wan, Ke
Li, Yucheng
Zhou, Yeyang
Chang, Li-Wen
Gu, Jiuxiang
Dong, Zhen
Anandkumar, Anima
Asi, Abedelkadir
Hu, Junjie
contents Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
Cai, Zefan
Xiao, Wen
Sun, Hanshi
Luo, Cheng
Zhang, Yikai
Wan, Ke
Li, Yucheng
Zhou, Yeyang
Chang, Li-Wen
Gu, Jiuxiang
Dong, Zhen
Anandkumar, Anima
Asi, Abedelkadir
Hu, Junjie
Computation and Language
Artificial Intelligence
Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.
title R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.24133