KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yang, Yifei, Cao, Zouying, Chen, Qiguang, Qin, Libo, Yang, Dongjie, Zhao, Hai, Chen, Zhi
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909362064195584
author Yang, Yifei
Cao, Zouying
Chen, Qiguang
Qin, Libo
Yang, Dongjie
Zhao, Hai
Chen, Zhi
author_facet Yang, Yifei
Cao, Zouying
Chen, Qiguang
Qin, Libo
Yang, Dongjie
Zhao, Hai
Chen, Zhi
contents The development of large language models (LLMs) has significantly expanded model sizes, resulting in substantial GPU memory requirements during inference. The key and value storage of the attention map in the KV (key-value) cache accounts for more than 80\% of this memory consumption. Nowadays, most existing KV cache compression methods focus on intra-layer compression within a single Transformer layer but few works consider layer-wise compression. In this paper, we propose a plug-and-play method called \textit{KVSharer}, which shares the KV cache between layers to achieve layer-wise compression. Rather than intuitively sharing based on higher similarity, we discover a counterintuitive phenomenon: sharing dissimilar KV caches better preserves the model performance. Experiments show that \textit{KVSharer} can reduce KV cache computation by 30\%, thereby lowering memory consumption without significantly impacting model performance and it can also achieve at least 1.3 times generation acceleration. Additionally, we verify that \textit{KVSharer} is compatible with existing intra-layer KV cache compression methods, and combining both can further save memory.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18517
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
Yang, Yifei
Cao, Zouying
Chen, Qiguang
Qin, Libo
Yang, Dongjie
Zhao, Hai
Chen, Zhi
Machine Learning
Artificial Intelligence
Computation and Language
The development of large language models (LLMs) has significantly expanded model sizes, resulting in substantial GPU memory requirements during inference. The key and value storage of the attention map in the KV (key-value) cache accounts for more than 80\% of this memory consumption. Nowadays, most existing KV cache compression methods focus on intra-layer compression within a single Transformer layer but few works consider layer-wise compression. In this paper, we propose a plug-and-play method called \textit{KVSharer}, which shares the KV cache between layers to achieve layer-wise compression. Rather than intuitively sharing based on higher similarity, we discover a counterintuitive phenomenon: sharing dissimilar KV caches better preserves the model performance. Experiments show that \textit{KVSharer} can reduce KV cache computation by 30\%, thereby lowering memory consumption without significantly impacting model performance and it can also achieve at least 1.3 times generation acceleration. Additionally, we verify that \textit{KVSharer} is compatible with existing intra-layer KV cache compression methods, and combining both can further save memory.
title KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.18517