On the Limits of Learned Importance Scoring for KV Cache Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Steele, Brady
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915742295785472
author Steele, Brady
author_facet Steele, Brady
contents We investigate learned KV cache compression through Speculative Importance Prediction (SIP), a 1.7M parameter non-query-aware scorer that predicts token importance from KV representations alone. Despite architectural sophistication (multi-horizon lookahead, cross-attention), SIP does not outperform simple baselines, including random selection, across 5 seeds, 4 retention levels, and 3 tasks. Key findings: (1) position-based heuristics (keep first 4 + last N tokens) match or exceed learned approaches; (2) prefill attention provides equivalent signal to complex learned scorers; (3) marginal information in KV representations beyond position and prefill attention appears limited for importance prediction. We hypothesize that circular dependence between future queries and generation trajectories contributes to this difficulty.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14279
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle On the Limits of Learned Importance Scoring for KV Cache Compression
Steele, Brady
Machine Learning
Artificial Intelligence
I.2.6; I.2.7; I.2.8
We investigate learned KV cache compression through Speculative Importance Prediction (SIP), a 1.7M parameter non-query-aware scorer that predicts token importance from KV representations alone. Despite architectural sophistication (multi-horizon lookahead, cross-attention), SIP does not outperform simple baselines, including random selection, across 5 seeds, 4 retention levels, and 3 tasks. Key findings: (1) position-based heuristics (keep first 4 + last N tokens) match or exceed learned approaches; (2) prefill attention provides equivalent signal to complex learned scorers; (3) marginal information in KV representations beyond position and prefill attention appears limited for importance prediction. We hypothesize that circular dependence between future queries and generation trajectories contributes to this difficulty.
title On the Limits of Learned Importance Scoring for KV Cache Compression
topic Machine Learning
Artificial Intelligence
I.2.6; I.2.7; I.2.8
url https://arxiv.org/abs/2601.14279