Draft-based Approximate Inference for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galim, Kevin, Ewer, Ethan, Kang, Wonjun, Lee, Minjae, Koo, Hyung Il, Lee, Kangwook
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910008079286272
author Galim, Kevin
Ewer, Ethan
Kang, Wonjun
Lee, Minjae
Koo, Hyung Il
Lee, Kangwook
author_facet Galim, Kevin
Ewer, Ethan
Kang, Wonjun
Lee, Minjae
Koo, Hyung Il
Lee, Kangwook
contents Optimizing inference for long-context large language models (LLMs) is increasingly important due to the quadratic compute and linear memory cost of Transformers. Existing approximate inference methods, including key-value (KV) cache dropping, sparse attention, and prompt compression, typically rely on coarse predictions of token or KV pair importance. We unify and extend recent work by introducing a framework for approximate LLM inference that leverages small draft models to more accurately predict token and KV pair importance. We provide novel theoretical and empirical analyses justifying lookahead-based importance estimation techniques. Within this framework, we present: (i) SpecKV, the first method to use lookahead with a small draft model to enable precise KV cache dropping; (ii) SpecPC, which leverages draft model attention activations to identify and discard less important prompt tokens; and (iii) SpecKV-PC, a cascaded compression strategy combining both techniques. Extensive experiments on long-context benchmarks demonstrate that our methods consistently achieve higher accuracy than existing baselines while retaining the same efficiency gains in memory usage, latency, and throughput.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Draft-based Approximate Inference for LLMs
Galim, Kevin
Ewer, Ethan
Kang, Wonjun
Lee, Minjae
Koo, Hyung Il
Lee, Kangwook
Computation and Language
Artificial Intelligence
Optimizing inference for long-context large language models (LLMs) is increasingly important due to the quadratic compute and linear memory cost of Transformers. Existing approximate inference methods, including key-value (KV) cache dropping, sparse attention, and prompt compression, typically rely on coarse predictions of token or KV pair importance. We unify and extend recent work by introducing a framework for approximate LLM inference that leverages small draft models to more accurately predict token and KV pair importance. We provide novel theoretical and empirical analyses justifying lookahead-based importance estimation techniques. Within this framework, we present: (i) SpecKV, the first method to use lookahead with a small draft model to enable precise KV cache dropping; (ii) SpecPC, which leverages draft model attention activations to identify and discard less important prompt tokens; and (iii) SpecKV-PC, a cascaded compression strategy combining both techniques. Extensive experiments on long-context benchmarks demonstrate that our methods consistently achieve higher accuracy than existing baselines while retaining the same efficiency gains in memory usage, latency, and throughput.
title Draft-based Approximate Inference for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2506.08373