DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dehghanighobadi, Zahra, Fischer, Asja
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915961064390656
author Dehghanighobadi, Zahra
Fischer, Asja
author_facet Dehghanighobadi, Zahra
Fischer, Asja
contents Long-context reasoning is a critical capability of large language models (LLMs), enabling applications such as long-document understanding, summarization, and code generation. However, efficient autoregressive inference relies on the key-value (KV) cache, whose memory footprint grows linearly with sequence length, leading to a major memory bottleneck. To mitigate this overhead, KV cache pruning methods discard cached tokens with low attention scores during inference. Most existing methods apply a uniform pruning ratio across layers, implicitly assuming that all layers contribute equally to overall model performance. We show that this assumption is suboptimal, as layers differ significantly in their sensitivity to pruning. We propose DepthKV, a layer-dependent pruning framework that allocates a fixed global KV budget across layers based on their sensitivity, rather than using a uniform allocation. Across multiple models and tasks, DepthKV consistently outperforms uniform pruning at the same global pruning ratio, demonstrating more effective utilization of the KV cache budget through layer-dependent allocation.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24647
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
Dehghanighobadi, Zahra
Fischer, Asja
Computation and Language
Artificial Intelligence
Long-context reasoning is a critical capability of large language models (LLMs), enabling applications such as long-document understanding, summarization, and code generation. However, efficient autoregressive inference relies on the key-value (KV) cache, whose memory footprint grows linearly with sequence length, leading to a major memory bottleneck. To mitigate this overhead, KV cache pruning methods discard cached tokens with low attention scores during inference. Most existing methods apply a uniform pruning ratio across layers, implicitly assuming that all layers contribute equally to overall model performance. We show that this assumption is suboptimal, as layers differ significantly in their sensitivity to pruning. We propose DepthKV, a layer-dependent pruning framework that allocates a fixed global KV budget across layers based on their sensitivity, rather than using a uniform allocation. Across multiple models and tasks, DepthKV consistently outperforms uniform pruning at the same global pruning ratio, demonstrating more effective utilization of the KV cache budget through layer-dependent allocation.
title DepthKV: Layer-Dependent KV Cache Pruning for Long-Context LLM Inference
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.24647