One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Liming, Qiu, Kaixi, Zhou, Jiayu, Kai, Jushi, Zhang, Haoyan, Wang, Huanyu, Leng, Jingwen, He, Ziwei, Lin, Zhouhan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915834972078080
author Lu, Liming
Qiu, Kaixi
Zhou, Jiayu
Kai, Jushi
Zhang, Haoyan
Wang, Huanyu
Leng, Jingwen
He, Ziwei
Lin, Zhouhan
author_facet Lu, Liming
Qiu, Kaixi
Zhou, Jiayu
Kai, Jushi
Zhang, Haoyan
Wang, Huanyu
Leng, Jingwen
He, Ziwei
Lin, Zhouhan
contents Despite the remarkable progress of Large Language Models (LLMs), the escalating memory footprint of the Key-Value (KV) cache remains a critical bottleneck for efficient inference. While dimensionality reduction offers a promising compression avenue, existing approaches typically either necessitate prohibitively expensive pre-training from scratch or suffer from severe performance deterioration under high compression regimes. In this work, we propose DynaKV, a novel post-training framework for low-rank KV cache compression. To the best of our knowledge, DynaKV is the first method to dynamically allocate compression rates to individual tokens according to their semantic meaning, which allows it to achieve better fidelity at aggressive compression ratios. Extensive experiments demonstrate that our method consistently outperforms existing state-of-the-art compression techniques, achieving significant memory reduction while maintaining competitive generation quality. Furthermore, our approach is orthogonal to sequence-level pruning methods. When integrated with SnapKV, DynaKV retains only 6% of the KV cache while maintaining 94% of the baseline performance on the LongBench benchmark.
format Preprint
id arxiv_https___arxiv_org_abs_2603_04411
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
Lu, Liming
Qiu, Kaixi
Zhou, Jiayu
Kai, Jushi
Zhang, Haoyan
Wang, Huanyu
Leng, Jingwen
He, Ziwei
Lin, Zhouhan
Computation and Language
Artificial Intelligence
Machine Learning
Despite the remarkable progress of Large Language Models (LLMs), the escalating memory footprint of the Key-Value (KV) cache remains a critical bottleneck for efficient inference. While dimensionality reduction offers a promising compression avenue, existing approaches typically either necessitate prohibitively expensive pre-training from scratch or suffer from severe performance deterioration under high compression regimes. In this work, we propose DynaKV, a novel post-training framework for low-rank KV cache compression. To the best of our knowledge, DynaKV is the first method to dynamically allocate compression rates to individual tokens according to their semantic meaning, which allows it to achieve better fidelity at aggressive compression ratios. Extensive experiments demonstrate that our method consistently outperforms existing state-of-the-art compression techniques, achieving significant memory reduction while maintaining competitive generation quality. Furthermore, our approach is orthogonal to sequence-level pruning methods. When integrated with SnapKV, DynaKV retains only 6% of the KV cache while maintaining 94% of the baseline performance on the LongBench benchmark.
title One Size Does Not Fit All: Token-Wise Adaptive Compression for KV Cache
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.04411