SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Mengjie, Song, William J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912548162371584
author Li, Mengjie
Song, William J.
author_facet Li, Mengjie
Song, William J.
contents The increasing input sequence length in Large Language Models (LLMs) puts significant pressure on key-value (KV) cache storage, making efficient inference challenging. Explicitly distinguishing attention behavior into our self-defined surface memorization and logic construction reveals essential roles in long-context reasoning. We observe that an individual attention head can display various behaviors, with nearly 98.5% effectively ignoring completely irrelevant information. The remaining 1.5% behaves as logic construction, and 0.5% behaves as surface memorization. Based on layer- and head-wise integration, we propose a novel two-stage SurfaceLogicKV method to utilize these attention behaviors for KV Cache compression. As a result, it achieves improved compressing robustness while maintaining competitive performance across various tasks and long sequences compared to baselines or even FullKV in some specific situations
format Preprint
id arxiv_https___arxiv_org_abs_2508_15806
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
Li, Mengjie
Song, William J.
Computation and Language
Artificial Intelligence
The increasing input sequence length in Large Language Models (LLMs) puts significant pressure on key-value (KV) cache storage, making efficient inference challenging. Explicitly distinguishing attention behavior into our self-defined surface memorization and logic construction reveals essential roles in long-context reasoning. We observe that an individual attention head can display various behaviors, with nearly 98.5% effectively ignoring completely irrelevant information. The remaining 1.5% behaves as logic construction, and 0.5% behaves as surface memorization. Based on layer- and head-wise integration, we propose a novel two-stage SurfaceLogicKV method to utilize these attention behaviors for KV Cache compression. As a result, it achieves improved compressing robustness while maintaining competitive performance across various tasks and long sequences compared to baselines or even FullKV in some specific situations
title SurfaceLogicKV: Surface and Logic Attention Behaviors are All You Need for Robust KV Cache Compression
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.15806