TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joshi, Vinay, Brahma, Pratik Prabhanjan, Liu, Zicheng, Barsoum, Emad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916780959596544
author Joshi, Vinay
Brahma, Pratik Prabhanjan
Liu, Zicheng
Barsoum, Emad
author_facet Joshi, Vinay
Brahma, Pratik Prabhanjan
Liu, Zicheng
Barsoum, Emad
contents The key-value (KV) cache in transformer models is a critical component for efficient decoding or inference, yet its memory demands scale poorly with sequence length, posing a major challenge for scalable deployment of large language models. Among several approaches to KV cache compression, quantization of key and value activations has been widely explored. Most KV cache quantization methods still need to manage sparse and noncontiguous outliers separately. To address this, we introduce TaDA, a training-free recipe for KV cache compression with quantization precision that adapts to error sensitivity across layers and a mean centering to eliminate separate outlier handling. Our approach yields substantial accuracy improvements for multiple models supporting various context lengths. Moreover, our approach does not need to separately manage outlier elements -- a persistent hurdle in most traditional quantization methods. Experiments on standard benchmarks demonstrate that our technique reduces KV cache memory footprint to 27% of the original 16-bit baseline while achieving comparable accuracy. Our method paves the way for scalable and high-performance reasoning in language models by potentially enabling inference for longer context length models, reasoning models, and longer chain of thoughts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_04642
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
Joshi, Vinay
Brahma, Pratik Prabhanjan
Liu, Zicheng
Barsoum, Emad
Computation and Language
The key-value (KV) cache in transformer models is a critical component for efficient decoding or inference, yet its memory demands scale poorly with sequence length, posing a major challenge for scalable deployment of large language models. Among several approaches to KV cache compression, quantization of key and value activations has been widely explored. Most KV cache quantization methods still need to manage sparse and noncontiguous outliers separately. To address this, we introduce TaDA, a training-free recipe for KV cache compression with quantization precision that adapts to error sensitivity across layers and a mean centering to eliminate separate outlier handling. Our approach yields substantial accuracy improvements for multiple models supporting various context lengths. Moreover, our approach does not need to separately manage outlier elements -- a persistent hurdle in most traditional quantization methods. Experiments on standard benchmarks demonstrate that our technique reduces KV cache memory footprint to 27% of the original 16-bit baseline while achieving comparable accuracy. Our method paves the way for scalable and high-performance reasoning in language models by potentially enabling inference for longer context length models, reasoning models, and longer chain of thoughts.
title TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
topic Computation and Language
url https://arxiv.org/abs/2506.04642