SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yi, Peng, Yajuan, Nguyen, Cam-Tu, Li, Zuchao, Wang, Xiaoliang, Zhao, Hai, Fu, Xiaoming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914262154215424
author Zhao, Yi
Peng, Yajuan
Nguyen, Cam-Tu
Li, Zuchao
Wang, Xiaoliang
Zhao, Hai
Fu, Xiaoming
author_facet Zhao, Yi
Peng, Yajuan
Nguyen, Cam-Tu
Li, Zuchao
Wang, Xiaoliang
Zhao, Hai
Fu, Xiaoming
contents KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens equally, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs of different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model's attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02751
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
Zhao, Yi
Peng, Yajuan
Nguyen, Cam-Tu
Li, Zuchao
Wang, Xiaoliang
Zhao, Hai
Fu, Xiaoming
Machine Learning
Artificial Intelligence
KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods often overlook two critical aspects: (1) their irreversible eviction strategy fails to adapt to dynamic attention patterns during decoding (the saliency shift problem), and (2) they treat both marginally important tokens and truly unimportant tokens equally, despite the collective significance of marginal tokens to model performance (the marginal information over-compression problem). To address these issues, we design two compensation mechanisms based on the high similarity of attention matrices between LLMs of different scales. We propose SmallKV, a small model assisted compensation method for KV cache compression. SmallKV can maintain attention matching between different-scale LLMs to: 1) assist the larger model in perceiving globally important information of attention; and 2) use the smaller model's attention scores to approximate those of marginal tokens in the larger model. Extensive experiments on benchmarks including GSM8K, BBH, MT-Bench, and LongBench demonstrate the effectiveness of SmallKV. Moreover, efficiency evaluations show that SmallKV achieves 1.75 - 2.56 times higher throughput than baseline methods, highlighting its potential for efficient and performant LLM inference in resource constrained environments.
title SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.02751