Lossless KV Cache Compression to 2%

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Zhen, Han, J. N., Wu, Kan, Xie, Ruobing, Wang, An, Sun, Xingwu, Kang, Zhanhui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916447494602752
author Yang, Zhen
Han, J. N.
Wu, Kan
Xie, Ruobing
Wang, An
Sun, Xingwu
Kang, Zhanhui
author_facet Yang, Zhen
Han, J. N.
Wu, Kan
Xie, Ruobing
Wang, An
Sun, Xingwu
Kang, Zhanhui
contents Large language models have revolutionized data processing in numerous domains, with their ability to handle extended context reasoning receiving notable recognition. To speed up inference, maintaining a key-value (KV) cache memory is essential. Nonetheless, the growing demands for KV cache memory create significant hurdles for efficient implementation. This work introduces a novel architecture, Cross-Layer Latent Attention (CLLA), aimed at compressing the KV cache to less than 2% of its original size while maintaining comparable performance levels. CLLA integrates multiple aspects of KV cache compression, including attention head/dimension reduction, layer sharing, and quantization techniques, into a cohesive framework. Our extensive experiments demonstrate that CLLA achieves lossless performance on most tasks while utilizing minimal KV cache, marking a significant advancement in practical KV cache compression.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15252
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lossless KV Cache Compression to 2%
Yang, Zhen
Han, J. N.
Wu, Kan
Xie, Ruobing
Wang, An
Sun, Xingwu
Kang, Zhanhui
Computation and Language
Artificial Intelligence
Large language models have revolutionized data processing in numerous domains, with their ability to handle extended context reasoning receiving notable recognition. To speed up inference, maintaining a key-value (KV) cache memory is essential. Nonetheless, the growing demands for KV cache memory create significant hurdles for efficient implementation. This work introduces a novel architecture, Cross-Layer Latent Attention (CLLA), aimed at compressing the KV cache to less than 2% of its original size while maintaining comparable performance levels. CLLA integrates multiple aspects of KV cache compression, including attention head/dimension reduction, layer sharing, and quantization techniques, into a cohesive framework. Our extensive experiments demonstrate that CLLA achieves lossless performance on most tasks while utilizing minimal KV cache, marking a significant advancement in practical KV cache compression.
title Lossless KV Cache Compression to 2%
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.15252