LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866910199205330944 |
|---|---|
| author | Zhou, Enshuai Hao, Yifan Wang, Chao Zhang, Rui Huang, Di Guo, Jiaming Hu, Xing Du, Zidong Guo, Qi Chen, Yunji |
| author_facet | Zhou, Enshuai Hao, Yifan Wang, Chao Zhang, Rui Huang, Di Guo, Jiaming Hu, Xing Du, Zidong Guo, Qi Chen, Yunji |
| contents | Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentally limited by heuristics: heuristic budgeting relies on statistical priors rather than task objectives, causing resource misallocation, while heuristic selection relies on coupled query-key interactions or static inductive biases (e.g., attention sinks). To address this limitation, we introduce LKV (Learned KV Eviction), which formulates KV compression as an end-to-end differentiable optimization problem. LKV integrates LKV-H to learn task-optimized global budgets, and LKV-T to derive intrinsic KV importance without materializing attention matrices. This design bypasses heuristic proxies, strictly aligning compression with task objectives. Extensive evaluations demonstrate that LKV achieves state-of-the-art performance on both LongBench and RULER benchmarks at high compression rates. In particular, on LongBench, LKV achieves near-lossless performance with only 15\% KV cache retention. Crucially, our analysis identifies learned budgeting as the dominant driver of fidelity, demonstrating that data-driven allocation is essential to overcome the limitations of hand-crafted heuristics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_06676 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction Zhou, Enshuai Hao, Yifan Wang, Chao Zhang, Rui Huang, Di Guo, Jiaming Hu, Xing Du, Zidong Guo, Qi Chen, Yunji Machine Learning Computation and Language Long-context inference in Large Language Models (LLMs) is bottlenecked by the linear growth of Key-Value (KV) cache memory. Existing KV cache compression paradigms are fundamentally limited by heuristics: heuristic budgeting relies on statistical priors rather than task objectives, causing resource misallocation, while heuristic selection relies on coupled query-key interactions or static inductive biases (e.g., attention sinks). To address this limitation, we introduce LKV (Learned KV Eviction), which formulates KV compression as an end-to-end differentiable optimization problem. LKV integrates LKV-H to learn task-optimized global budgets, and LKV-T to derive intrinsic KV importance without materializing attention matrices. This design bypasses heuristic proxies, strictly aligning compression with task objectives. Extensive evaluations demonstrate that LKV achieves state-of-the-art performance on both LongBench and RULER benchmarks at high compression rates. In particular, on LongBench, LKV achieves near-lossless performance with only 15\% KV cache retention. Crucially, our analysis identifies learned budgeting as the dominant driver of fidelity, demonstrating that data-driven allocation is essential to overcome the limitations of hand-crafted heuristics. |
| title | LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2605.06676 |