CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fei, Chao, Li, Guozhong, Liu, Chenxi, Kalnis, Panos
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910031420588032
author Fei, Chao
Li, Guozhong
Liu, Chenxi
Kalnis, Panos
author_facet Fei, Chao
Li, Guozhong
Liu, Chenxi
Kalnis, Panos
contents Long-context LLMs demand accurate inference at low latency, yet decoding becomes primarily constrained by KV cache as context grows. Prior pruning methods are largely context-agnostic: their token selection ignores step-wise relevance and local semantics, which undermines quality. Moreover, their irregular accesses and selection overheads yield only limited wall-clock speedups. To address this, we propose \textbf{CHESS}, an \textit{algorithm-system co-design} KV-cache management system. Algorithmically, CHESS introduces a context-aware, hierarchical selection policy that dynamically reconstructs a coherent context for the current decoding. System-wise, coarse granularity selection eliminates expensive data movement, fully realizing practical acceleration from theoretical sparsity. Extensive evaluations demonstrate that CHESS surpasses Full-KV quality using only \textbf{1\%} of the KV cache, delivers low-latency stable inference with up to \textbf{4.56$\times$} higher throughput, and consistently outperforms other strong baselines. Code is available at \href{https://anonymous.4open.science/r/CHESS-9958/}{https://anonymous.4open.science/r/CHESS/}.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20732
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference
Fei, Chao
Li, Guozhong
Liu, Chenxi
Kalnis, Panos
Artificial Intelligence
Long-context LLMs demand accurate inference at low latency, yet decoding becomes primarily constrained by KV cache as context grows. Prior pruning methods are largely context-agnostic: their token selection ignores step-wise relevance and local semantics, which undermines quality. Moreover, their irregular accesses and selection overheads yield only limited wall-clock speedups. To address this, we propose \textbf{CHESS}, an \textit{algorithm-system co-design} KV-cache management system. Algorithmically, CHESS introduces a context-aware, hierarchical selection policy that dynamically reconstructs a coherent context for the current decoding. System-wise, coarse granularity selection eliminates expensive data movement, fully realizing practical acceleration from theoretical sparsity. Extensive evaluations demonstrate that CHESS surpasses Full-KV quality using only \textbf{1\%} of the KV cache, delivers low-latency stable inference with up to \textbf{4.56$\times$} higher throughput, and consistently outperforms other strong baselines. Code is available at \href{https://anonymous.4open.science/r/CHESS-9958/}{https://anonymous.4open.science/r/CHESS/}.
title CHESS: Context-aware Hierarchical Efficient Semantic Selection for Long-Context LLM Inference
topic Artificial Intelligence
url https://arxiv.org/abs/2602.20732