RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tang, Hanlin, Lin, Yang, Lin, Jing, Han, Qingsen, Hong, Shikuan, Yao, Yiwu, Wang, Gongyi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913441623572480
author Tang, Hanlin
Lin, Yang
Lin, Jing
Han, Qingsen
Hong, Shikuan
Yao, Yiwu
Wang, Gongyi
author_facet Tang, Hanlin
Lin, Yang
Lin, Jing
Han, Qingsen
Hong, Shikuan
Yao, Yiwu
Wang, Gongyi
contents The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, which irreversibly erases critical information that might be needed for future queries. In this paper, we propose a novel compression technique for KV cache that preserves all token information. Our investigation reveals that: i) Most attention heads primarily focus on the local context; ii) Only a few heads, denoted as retrieval heads, can essentially pay attention to all input tokens. These key observations motivate us to use separate caching strategy for attention heads. Therefore, we propose RazorAttention, a training-free KV cache compression algorithm, which maintains a full cache for these crucial retrieval heads and discards the remote tokens in non-retrieval heads. Furthermore, we introduce a novel mechanism involving a "compensation token" to further recover the information in the dropped tokens. Extensive evaluations across a diverse set of large language models (LLMs) demonstrate that RazorAttention achieves a reduction in KV cache size by over 70% without noticeable impacts on performance. Additionally, RazorAttention is compatible with FlashAttention, rendering it an efficient and plug-and-play solution that enhances LLM inference efficiency without overhead or retraining of the original model.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
Tang, Hanlin
Lin, Yang
Lin, Jing
Han, Qingsen
Hong, Shikuan
Yao, Yiwu
Wang, Gongyi
Machine Learning
Computation and Language
The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, which irreversibly erases critical information that might be needed for future queries. In this paper, we propose a novel compression technique for KV cache that preserves all token information. Our investigation reveals that: i) Most attention heads primarily focus on the local context; ii) Only a few heads, denoted as retrieval heads, can essentially pay attention to all input tokens. These key observations motivate us to use separate caching strategy for attention heads. Therefore, we propose RazorAttention, a training-free KV cache compression algorithm, which maintains a full cache for these crucial retrieval heads and discards the remote tokens in non-retrieval heads. Furthermore, we introduce a novel mechanism involving a "compensation token" to further recover the information in the dropped tokens. Extensive evaluations across a diverse set of large language models (LLMs) demonstrate that RazorAttention achieves a reduction in KV cache size by over 70% without noticeable impacts on performance. Additionally, RazorAttention is compatible with FlashAttention, rendering it an efficient and plug-and-play solution that enhances LLM inference efficiency without overhead or retraining of the original model.
title RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2407.15891