KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Tianyi, Yi, Jonah, Xu, Zhaozhuo, Shrivastava, Anshumali
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910436681580544
author Zhang, Tianyi
Yi, Jonah
Xu, Zhaozhuo
Shrivastava, Anshumali
author_facet Zhang, Tianyi
Yi, Jonah
Xu, Zhaozhuo
Shrivastava, Anshumali
contents Efficient deployment of Large Language Models (LLMs) requires batching multiple requests together to improve throughput. As the batch size, context length, or model size increases, the size of the key and value (KV) cache can quickly become the main contributor to GPU memory usage and the bottleneck of inference latency. Quantization has emerged as an effective technique for KV cache compression, but existing methods still fail at very low bit widths. We observe that distinct channels of a key/value activation embedding are highly inter-dependent, and the joint entropy of multiple channels grows at a slower rate than the sum of their marginal entropies. Based on this insight, we propose Coupled Quantization (CQ), which couples multiple key/value channels together to exploit their inter-dependency and encode the activations in a more information-efficient manner. Extensive experiments reveal that CQ outperforms or is competitive with existing baselines in preserving model quality. Furthermore, we demonstrate that CQ can preserve model quality with KV cache quantized down to 1-bit.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03917
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
Zhang, Tianyi
Yi, Jonah
Xu, Zhaozhuo
Shrivastava, Anshumali
Machine Learning
Efficient deployment of Large Language Models (LLMs) requires batching multiple requests together to improve throughput. As the batch size, context length, or model size increases, the size of the key and value (KV) cache can quickly become the main contributor to GPU memory usage and the bottleneck of inference latency. Quantization has emerged as an effective technique for KV cache compression, but existing methods still fail at very low bit widths. We observe that distinct channels of a key/value activation embedding are highly inter-dependent, and the joint entropy of multiple channels grows at a slower rate than the sum of their marginal entropies. Based on this insight, we propose Coupled Quantization (CQ), which couples multiple key/value channels together to exploit their inter-dependency and encode the activations in a more information-efficient manner. Extensive experiments reveal that CQ outperforms or is competitive with existing baselines in preserving model quality. Furthermore, we demonstrate that CQ can preserve model quality with KV cache quantized down to 1-bit.
title KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
topic Machine Learning
url https://arxiv.org/abs/2405.03917