LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhou, Enshuai, Hao, Yifan, Wang, Chao, Zhang, Rui, Huang, Di, Guo, Jiaming, Hu, Xing, Du, Zidong, Guo, Qi, Chen, Yunji
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910199205330944
author Zhou, Enshuai
Hao, Yifan
Wang, Chao
Zhang, Rui
Huang, Di
Guo, Jiaming
Hu, Xing
Du, Zidong
Guo, Qi
Chen, Yunji
author_facet Zhou, Enshuai
Hao, Yifan
Wang, Chao
Zhang, Rui
Huang, Di
Guo, Jiaming
Hu, Xing
Du, Zidong
Guo, Qi
Chen, Yunji
contents Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentally limited by heuristics: heuristic budgeting relies on statistical priors rather than task objectives, causing resource misallocation, while heuristic selection relies on coupled query-key interactions or static inductive biases (e.g., attention sinks). To address this limitation, we introduce LKV (Learned KV Eviction), which formulates KV compression as an end-to-end differentiable optimization problem. LKV integrates LKV-H to learn task-optimized global budgets, and LKV-T to derive intrinsic KV importance without materializing attention matrices. This design bypasses heuristic proxies, strictly aligning compression with task objectives. Extensive evaluations demonstrate that LKV achieves state-of-the-art performance on both LongBench and RULER benchmarks at high compression rates. In particular, on LongBench, LKV achieves near-lossless performance with only 15\% KV cache retention. Crucially, our analysis identifies learned budgeting as the dominant driver of fidelity, demonstrating that data-driven allocation is essential to overcome the limitations of hand-crafted heuristics.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06676
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
Zhou, Enshuai
Hao, Yifan
Wang, Chao
Zhang, Rui
Huang, Di
Guo, Jiaming
Hu, Xing
Du, Zidong
Guo, Qi
Chen, Yunji
Machine Learning
Computation and Language
Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentally limited by heuristics: heuristic budgeting relies on statistical priors rather than task objectives, causing resource misallocation, while heuristic selection relies on coupled query-key interactions or static inductive biases (e.g., attention sinks). To address this limitation, we introduce LKV (Learned KV Eviction), which formulates KV compression as an end-to-end differentiable optimization problem. LKV integrates LKV-H to learn task-optimized global budgets, and LKV-T to derive intrinsic KV importance without materializing attention matrices. This design bypasses heuristic proxies, strictly aligning compression with task objectives. Extensive evaluations demonstrate that LKV achieves state-of-the-art performance on both LongBench and RULER benchmarks at high compression rates. In particular, on LongBench, LKV achieves near-lossless performance with only 15\% KV cache retention. Crucially, our analysis identifies learned budgeting as the dominant driver of fidelity, demonstrating that data-driven allocation is essential to overcome the limitations of hand-crafted heuristics.
title LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.06676