Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916701062299648 |
|---|---|
| author | Javidnia, Neusha Rouhani, Bita Darvish Koushanfar, Farinaz |
| author_facet | Javidnia, Neusha Rouhani, Bita Darvish Koushanfar, Farinaz |
| contents | Large language models (LLMs) have demonstrated exceptional capabilities in generating text, images, and video content. However, as context length grows, the computational cost of attention increases quadratically with the number of tokens, presenting significant efficiency challenges. This paper presents an analysis of various Key-Value (KV) cache compression strategies, offering a comprehensive taxonomy that categorizes these methods by their underlying principles and implementation techniques. Furthermore, we evaluate their impact on performance and inference latency, providing critical insights into their effectiveness. Our findings highlight the trade-offs involved in KV cache compression and its influence on handling long-context scenarios, paving the way for more efficient LLM implementations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_11816 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques Javidnia, Neusha Rouhani, Bita Darvish Koushanfar, Farinaz Computation and Language Large language models (LLMs) have demonstrated exceptional capabilities in generating text, images, and video content. However, as context length grows, the computational cost of attention increases quadratically with the number of tokens, presenting significant efficiency challenges. This paper presents an analysis of various Key-Value (KV) cache compression strategies, offering a comprehensive taxonomy that categorizes these methods by their underlying principles and implementation techniques. Furthermore, we evaluate their impact on performance and inference latency, providing critical insights into their effectiveness. Our findings highlight the trade-offs involved in KV cache compression and its influence on handling long-context scenarios, paving the way for more efficient LLM implementations. |
| title | Key, Value, Compress: A Systematic Exploration of KV Cache Compression Techniques |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2503.11816 |