PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Dongjie, Han, XiaoDong, Gao, Yan, Hu, Yao, Zhang, Shilin, Zhao, Hai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929374249353216
author Yang, Dongjie
Han, XiaoDong
Gao, Yan
Hu, Yao
Zhang, Shilin
Zhao, Hai
author_facet Yang, Dongjie
Han, XiaoDong
Gao, Yan
Hu, Yao
Zhang, Shilin
Zhao, Hai
contents Large Language Models (LLMs) have shown remarkable comprehension abilities but face challenges in GPU memory usage during inference, hindering their scalability for real-time applications like chatbots. To accelerate inference, we store computed keys and values (KV cache) in the GPU memory. Existing methods study the KV cache compression to reduce memory by pruning the pre-computed KV cache. However, they neglect the inter-layer dependency between layers and huge memory consumption in pre-computation. To explore these deficiencies, we find that the number of crucial keys and values that influence future generations decreases layer by layer and we can extract them by the consistency in attention weights. Based on the findings, we propose PyramidInfer, a method that compresses the KV cache by layer-wise retaining crucial context. PyramidInfer saves significant memory by computing fewer keys and values without sacrificing performance. Experimental results show PyramidInfer improves 2.2x throughput compared to Accelerate with over 54% GPU memory reduction in KV cache.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12532
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Yang, Dongjie
Han, XiaoDong
Gao, Yan
Hu, Yao
Zhang, Shilin
Zhao, Hai
Computation and Language
Large Language Models (LLMs) have shown remarkable comprehension abilities but face challenges in GPU memory usage during inference, hindering their scalability for real-time applications like chatbots. To accelerate inference, we store computed keys and values (KV cache) in the GPU memory. Existing methods study the KV cache compression to reduce memory by pruning the pre-computed KV cache. However, they neglect the inter-layer dependency between layers and huge memory consumption in pre-computation. To explore these deficiencies, we find that the number of crucial keys and values that influence future generations decreases layer by layer and we can extract them by the consistency in attention weights. Based on the findings, we propose PyramidInfer, a method that compresses the KV cache by layer-wise retaining crucial context. PyramidInfer saves significant memory by computing fewer keys and values without sacrificing performance. Experimental results show PyramidInfer improves 2.2x throughput compared to Accelerate with over 54% GPU memory reduction in KV cache.
title PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
topic Computation and Language
url https://arxiv.org/abs/2405.12532