R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866911390815485952 |
|---|---|
| author | Cai, Zefan Xiao, Wen Sun, Hanshi Luo, Cheng Zhang, Yikai Wan, Ke Li, Yucheng Zhou, Yeyang Chang, Li-Wen Gu, Jiuxiang Dong, Zhen Anandkumar, Anima Asi, Abedelkadir Hu, Junjie |
| author_facet | Cai, Zefan Xiao, Wen Sun, Hanshi Luo, Cheng Zhang, Yikai Wan, Ke Li, Yucheng Zhou, Yeyang Chang, Li-Wen Gu, Jiuxiang Dong, Zhen Anandkumar, Anima Asi, Abedelkadir Hu, Junjie |
| contents | Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_24133 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | R-KV: Redundancy-aware KV Cache Compression for Reasoning Models Cai, Zefan Xiao, Wen Sun, Hanshi Luo, Cheng Zhang, Yikai Wan, Ke Li, Yucheng Zhou, Yeyang Chang, Li-Wen Gu, Jiuxiang Dong, Zhen Anandkumar, Anima Asi, Abedelkadir Hu, Junjie Computation and Language Artificial Intelligence Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reach only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6X throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets. |
| title | R-KV: Redundancy-aware KV Cache Compression for Reasoning Models |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2505.24133 |