KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gokhale, Sai, Das, Devleena, Patwari, Rajeev, Sirasao, Ashish, Delaye, Elliott
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917116499722240
author Gokhale, Sai
Das, Devleena
Patwari, Rajeev
Sirasao, Ashish
Delaye, Elliott
author_facet Gokhale, Sai
Das, Devleena
Patwari, Rajeev
Sirasao, Ashish
Delaye, Elliott
contents Long-context Large Language Models (LLMs) face significant memory bottlenecks during inference due to the linear growth of key-value (KV) cache with sequence length. While individual optimization techniques like KV cache quantization, chunked prefill, and model weight quantization have shown promise, their joint effects and optimal configurations for edge deployment remain underexplored. We introduce KV Pareto, a systems-level framework that systematically maps the trade-off frontier between total memory consumption and task accuracy across these three complementary optimization techniques. Our framework evaluates multiple LLM architectures (Qwen, Llama, Mistral) with varying KV quantization schemes (int2/4/8, mixed-precision), granularities (per-token, per-tensor, per-block), and 4-bit weight quantization via AWQ. Our framework identifies model-specific Pareto-optimal configurations that achieve 68-78% total memory reduction with minimal (1-3%) accuracy degradation on long-context tasks. We additionally verify the selected frontiers on additional benchmarks of Needle-in-a-Haystack, GSM8k and MMLU as well as extended context lengths of up to 128k to demonstrate the practical need of joint optimization for efficient LLM inference.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01953
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
Gokhale, Sai
Das, Devleena
Patwari, Rajeev
Sirasao, Ashish
Delaye, Elliott
Machine Learning
Long-context Large Language Models (LLMs) face significant memory bottlenecks during inference due to the linear growth of key-value (KV) cache with sequence length. While individual optimization techniques like KV cache quantization, chunked prefill, and model weight quantization have shown promise, their joint effects and optimal configurations for edge deployment remain underexplored. We introduce KV Pareto, a systems-level framework that systematically maps the trade-off frontier between total memory consumption and task accuracy across these three complementary optimization techniques. Our framework evaluates multiple LLM architectures (Qwen, Llama, Mistral) with varying KV quantization schemes (int2/4/8, mixed-precision), granularities (per-token, per-tensor, per-block), and 4-bit weight quantization via AWQ. Our framework identifies model-specific Pareto-optimal configurations that achieve 68-78% total memory reduction with minimal (1-3%) accuracy degradation on long-context tasks. We additionally verify the selected frontiers on additional benchmarks of Needle-in-a-Haystack, GSM8k and MMLU as well as extended context lengths of up to 128k to demonstrate the practical need of joint optimization for efficient LLM inference.
title KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
topic Machine Learning
url https://arxiv.org/abs/2512.01953