Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jang-Hyun, Han, Dongyoon, Yun, Sangdoo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914315301289984
author Kim, Jang-Hyun
Han, Dongyoon
Yun, Sangdoo
author_facet Kim, Jang-Hyun
Han, Dongyoon
Yun, Sangdoo
contents Efficient key-value (KV) cache management is crucial for the practical deployment of large language models (LLMs), yet existing compression techniques often incur a trade-off between performance degradation and computational overhead. We propose a novel gating-based KV cache eviction method for frozen-weight LLMs that achieves high compression ratios with negligible computational cost. Our approach introduces lightweight sink-attention gating modules to identify and retain critical KV pairs, and integrates seamlessly into both the prefill and decoding stages. The proposed gate training algorithm relies on forward passes of an LLM, avoiding expensive backpropagation, while achieving strong task generalization through a task-agnostic reconstruction objective. Extensive experiments across the Qwen2.5-1M, Qwen3, and Gemma3 families show that our method maintains near-lossless performance while evicting up to 70% of the KV cache. The results are consistent across a wide range of tasks, including long-context understanding, code comprehension, and mathematical reasoning, demonstrating the generality of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17668
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
Kim, Jang-Hyun
Han, Dongyoon
Yun, Sangdoo
Machine Learning
Computation and Language
Efficient key-value (KV) cache management is crucial for the practical deployment of large language models (LLMs), yet existing compression techniques often incur a trade-off between performance degradation and computational overhead. We propose a novel gating-based KV cache eviction method for frozen-weight LLMs that achieves high compression ratios with negligible computational cost. Our approach introduces lightweight sink-attention gating modules to identify and retain critical KV pairs, and integrates seamlessly into both the prefill and decoding stages. The proposed gate training algorithm relies on forward passes of an LLM, avoiding expensive backpropagation, while achieving strong task generalization through a task-agnostic reconstruction objective. Extensive experiments across the Qwen2.5-1M, Qwen3, and Gemma3 families show that our method maintains near-lossless performance while evicting up to 70% of the KV cache. The results are consistent across a wide range of tasks, including long-context understanding, code comprehension, and mathematical reasoning, demonstrating the generality of our approach.
title Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2601.17668