KVzap: Fast, Adaptive, and Faithful KV Cache Pruning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jegou, Simon, Jeblick, Maximilian
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911419243429888
author Jegou, Simon
Jeblick, Maximilian
author_facet Jegou, Simon
Jeblick, Maximilian
contents Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation of KVzip that works in both prefilling and decoding. On Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B across long-context and reasoning tasks, KVzap achieves $2$--$4\times$ KV cache compression with negligible accuracy loss and achieves state-of-the-art performance on the KVpress leaderboard. Code and models are available at https://github.com/NVIDIA/kvpress.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07891
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
Jegou, Simon
Jeblick, Maximilian
Machine Learning
Artificial Intelligence
Computation and Language
Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation of KVzip that works in both prefilling and decoding. On Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B across long-context and reasoning tasks, KVzap achieves $2$--$4\times$ KV cache compression with negligible accuracy loss and achieves state-of-the-art performance on the KVpress leaderboard. Code and models are available at https://github.com/NVIDIA/kvpress.
title KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.07891