Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shi, Luohe, Zhang, Hongyi, Yao, Yao, Li, Zuchao, Zhao, Hai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913581395607552
author Shi, Luohe
Zhang, Hongyi
Yao, Yao
Li, Zuchao
Zhao, Hai
author_facet Shi, Luohe
Zhang, Hongyi
Yao, Yao
Li, Zuchao
Zhao, Hai
contents Large Language Models (LLMs), epitomized by ChatGPT's release in late 2022, have revolutionized various industries with their advanced language comprehension. However, their efficiency is challenged by the Transformer architecture's struggle with handling long texts. KV Cache has emerged as a pivotal solution to this issue, converting the time complexity of token generation from quadratic to linear, albeit with increased GPU memory overhead proportional to conversation length. With the development of the LLM community and academia, various KV Cache compression methods have been proposed. In this review, we dissect the various properties of KV Cache and elaborate on various methods currently used to optimize the KV Cache space usage of LLMs. These methods span the pre-training phase, deployment phase, and inference phase, and we summarize the commonalities and differences among these methods. Additionally, we list some metrics for evaluating the long-text capabilities of large language models, from both efficiency and capability perspectives. Our review thus sheds light on the evolving landscape of LLM optimization, offering insights into future advancements in this dynamic field. Links to the papers mentioned in this review can be found in our Github Repo https://github.com/zcli-charlie/Awesome-KV-Cache.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18003
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
Shi, Luohe
Zhang, Hongyi
Yao, Yao
Li, Zuchao
Zhao, Hai
Computation and Language
Large Language Models (LLMs), epitomized by ChatGPT's release in late 2022, have revolutionized various industries with their advanced language comprehension. However, their efficiency is challenged by the Transformer architecture's struggle with handling long texts. KV Cache has emerged as a pivotal solution to this issue, converting the time complexity of token generation from quadratic to linear, albeit with increased GPU memory overhead proportional to conversation length. With the development of the LLM community and academia, various KV Cache compression methods have been proposed. In this review, we dissect the various properties of KV Cache and elaborate on various methods currently used to optimize the KV Cache space usage of LLMs. These methods span the pre-training phase, deployment phase, and inference phase, and we summarize the commonalities and differences among these methods. Additionally, we list some metrics for evaluating the long-text capabilities of large language models, from both efficiency and capability perspectives. Our review thus sheds light on the evolving landscape of LLM optimization, offering insights into future advancements in this dynamic field. Links to the papers mentioned in this review can be found in our Github Repo https://github.com/zcli-charlie/Awesome-KV-Cache.
title Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
topic Computation and Language
url https://arxiv.org/abs/2407.18003