Saved in:
Bibliographic Details
Main Authors: Zeng, Hui, Zhao, Daming, Yang, Pengfei, Hou, WenXuan, Zheng, Tianyang, Li, Hui, Ji, Weiye, Zhai, Jidong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.06029
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908710904791040
author Zeng, Hui
Zhao, Daming
Yang, Pengfei
Hou, WenXuan
Zheng, Tianyang
Li, Hui
Ji, Weiye
Zhai, Jidong
author_facet Zeng, Hui
Zhao, Daming
Yang, Pengfei
Hou, WenXuan
Zheng, Tianyang
Li, Hui
Ji, Weiye
Zhai, Jidong
contents Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV) caches. While existing KV compression methods primarily focus on reducing prefill memory from long input sequences, they fall short in addressing the dynamic and layer-sensitive nature of long-form generation, which is central to reasoning tasks. We propose Lethe, a dynamic KV cache management framework that introduces adaptivity along both the spatial and temporal dimensions of decoding. Along the spatial dimension, Lethe performs layerwise sparsity-aware allocation, assigning token pruning budgets to each transformer layer based on estimated attention redundancy. Along the temporal dimension, Lethe conducts multi-round token pruning during generation, driven by a Recency-Aware Selective Retention} (RASR) mechanism. RASR extends traditional recency-based heuristics by also considering token relevance derived from evolving attention patterns, enabling informed decisions about which tokens to retain or evict. Empirical results demonstrate that Lethe achieves a favorable balance between efficiency and generation quality across diverse models and tasks, increases throughput by up to 2.56x.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06029
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
Zeng, Hui
Zhao, Daming
Yang, Pengfei
Hou, WenXuan
Zheng, Tianyang
Li, Hui
Ji, Weiye
Zhai, Jidong
Machine Learning
Generative reasoning with large language models (LLMs) often involves long decoding sequences, leading to substantial memory and latency overheads from accumulating key-value (KV) caches. While existing KV compression methods primarily focus on reducing prefill memory from long input sequences, they fall short in addressing the dynamic and layer-sensitive nature of long-form generation, which is central to reasoning tasks. We propose Lethe, a dynamic KV cache management framework that introduces adaptivity along both the spatial and temporal dimensions of decoding. Along the spatial dimension, Lethe performs layerwise sparsity-aware allocation, assigning token pruning budgets to each transformer layer based on estimated attention redundancy. Along the temporal dimension, Lethe conducts multi-round token pruning during generation, driven by a Recency-Aware Selective Retention} (RASR) mechanism. RASR extends traditional recency-based heuristics by also considering token relevance derived from evolving attention patterns, enabling informed decisions about which tokens to retain or evict. Empirical results demonstrate that Lethe achieves a favorable balance between efficiency and generation quality across diverse models and tasks, increases throughput by up to 2.56x.
title Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
topic Machine Learning
url https://arxiv.org/abs/2511.06029