Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Javidnia, Neusha, Rouhani, Bita Darvish, Koushanfar, Farinaz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916701062299648
author Javidnia, Neusha
Rouhani, Bita Darvish
Koushanfar, Farinaz
author_facet Javidnia, Neusha
Rouhani, Bita Darvish
Koushanfar, Farinaz
contents Large language models (LLMs) have demonstrated exceptional capabilities in generating text, images, and video content. However, as context length grows, the computational cost of attention increases quadratically with the number of tokens, presenting significant efficiency challenges. This paper presents an analysis of various Key-Value (KV) cache compression strategies, offering a comprehensive taxonomy that categorizes these methods by their underlying principles and implementation techniques. Furthermore, we evaluate their impact on performance and inference latency, providing critical insights into their effectiveness. Our findings highlight the trade-offs involved in KV cache compression and its influence on handling long-context scenarios, paving the way for more efficient LLM implementations.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11816
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
Javidnia, Neusha
Rouhani, Bita Darvish
Koushanfar, Farinaz
Computation and Language
Large language models (LLMs) have demonstrated exceptional capabilities in generating text, images, and video content. However, as context length grows, the computational cost of attention increases quadratically with the number of tokens, presenting significant efficiency challenges. This paper presents an analysis of various Key-Value (KV) cache compression strategies, offering a comprehensive taxonomy that categorizes these methods by their underlying principles and implementation techniques. Furthermore, we evaluate their impact on performance and inference latency, providing critical insights into their effectiveness. Our findings highlight the trade-offs involved in KV cache compression and its influence on handling long-context scenarios, paving the way for more efficient LLM implementations.
title Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
topic Computation and Language
url https://arxiv.org/abs/2503.11816