G-KV: Decoding-Time KV Cache Eviction with Global Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Mengqi, Wang, Lu, Zhang, Chaoyun, Shen, Zekai, Mao, Xiaowei, Qin, Si, Lin, Qingwei, Rajmohan, Saravan, Zhang, Dongmei, Wan, Huaiyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918224433512448
author Liao, Mengqi
Wang, Lu
Zhang, Chaoyun
Shen, Zekai
Mao, Xiaowei
Qin, Si
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Wan, Huaiyu
author_facet Liao, Mengqi
Wang, Lu
Zhang, Chaoyun
Shen, Zekai
Mao, Xiaowei
Qin, Si
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Wan, Huaiyu
contents Recent reasoning large language models (LLMs) excel in complex tasks but encounter significant computational and memory challenges due to long sequence lengths. KV cache compression has emerged as an effective approach to greatly enhance the efficiency of reasoning. However, existing methods often focus on prompt compression or token eviction with local attention score, overlooking the long-term importance of tokens. We propose G-KV, a KV cache eviction method that employs a global scoring mechanism, combining local and historical attention scores to more accurately assess token importance. Additionally, we introduce post-training techniques, including reinforcement learning and distillation, to optimize models for compressed KV cache settings. The code of this paper is available on: https://github.com/microsoft/G-KV.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00504
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle G-KV: Decoding-Time KV Cache Eviction with Global Attention
Liao, Mengqi
Wang, Lu
Zhang, Chaoyun
Shen, Zekai
Mao, Xiaowei
Qin, Si
Lin, Qingwei
Rajmohan, Saravan
Zhang, Dongmei
Wan, Huaiyu
Computation and Language
Artificial Intelligence
Recent reasoning large language models (LLMs) excel in complex tasks but encounter significant computational and memory challenges due to long sequence lengths. KV cache compression has emerged as an effective approach to greatly enhance the efficiency of reasoning. However, existing methods often focus on prompt compression or token eviction with local attention score, overlooking the long-term importance of tokens. We propose G-KV, a KV cache eviction method that employs a global scoring mechanism, combining local and historical attention scores to more accurately assess token importance. Additionally, we introduce post-training techniques, including reinforcement learning and distillation, to optimize models for compressed KV cache settings. The code of this paper is available on: https://github.com/microsoft/G-KV.
title G-KV: Decoding-Time KV Cache Eviction with Global Attention
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.00504