Inference Time Context Sparsity: Illusion or Opportunity?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Joshi, Sahil, Dixit, Prithvi, Chowdhury, Agniva, Shrivastava, Anshumali, Gonzalez, Joseph E., Stoica, Ion, Agrawal, Kumar Krishna, Desai, Aditya
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914594195243008
author Joshi, Sahil
Dixit, Prithvi
Chowdhury, Agniva
Shrivastava, Anshumali
Gonzalez, Joseph E.
Stoica, Ion
Agrawal, Kumar Krishna
Desai, Aditya
author_facet Joshi, Sahil
Dixit, Prithvi
Chowdhury, Agniva
Shrivastava, Anshumali
Gonzalez, Joseph E.
Stoica, Ion
Agrawal, Kumar Krishna
Desai, Aditya
contents Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position is that these constraints are artificial and unnecessary, and that the future of LLM inference lies in extreme but principled sparsity along the context dimension. This position is supported by several strands of empirical and theoretical evidence. First, we find the insistence on dense attention unreasonable, since in a long context a query effectively projects O(N) attention information into a hidden space of dimension d << N, making the process inherently lossy. Second, we perform an extensive study of sparsity in LLMs spanning 20 models across five model families, varying context lengths, and different sparsity levels. We empirically demonstrate a strong trend: current LLMs, despite not being trained for context sparsity, are remarkably robust to inference-time decode sparsity across tasks of varying complexity, including retrieval, multi-hop QA, mathematical reasoning, and agentic coding. Importantly, we also show that current hardware is already sufficient to realize substantial gains from this sparsity. For example, our sparse decode kernels accelerate large-context processing by up to 10x over FlashInfer at 50x sparsity levels on hardware such as the H100. Overall, these results position extreme context sparsity not as a heuristic, but as a principled foundation for LLM inference, training, and architecture design: one that is both feasible and beneficial, and a compelling direction for future systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_24168
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Inference Time Context Sparsity: Illusion or Opportunity?
Joshi, Sahil
Dixit, Prithvi
Chowdhury, Agniva
Shrivastava, Anshumali
Gonzalez, Joseph E.
Stoica, Ion
Agrawal, Kumar Krishna
Desai, Aditya
Artificial Intelligence
Machine Learning
Sparsity has long been a central theme in LLM efficiency, but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position is that these constraints are artificial and unnecessary, and that the future of LLM inference lies in extreme but principled sparsity along the context dimension. This position is supported by several strands of empirical and theoretical evidence. First, we find the insistence on dense attention unreasonable, since in a long context a query effectively projects O(N) attention information into a hidden space of dimension d << N, making the process inherently lossy. Second, we perform an extensive study of sparsity in LLMs spanning 20 models across five model families, varying context lengths, and different sparsity levels. We empirically demonstrate a strong trend: current LLMs, despite not being trained for context sparsity, are remarkably robust to inference-time decode sparsity across tasks of varying complexity, including retrieval, multi-hop QA, mathematical reasoning, and agentic coding. Importantly, we also show that current hardware is already sufficient to realize substantial gains from this sparsity. For example, our sparse decode kernels accelerate large-context processing by up to 10x over FlashInfer at 50x sparsity levels on hardware such as the H100. Overall, these results position extreme context sparsity not as a heuristic, but as a principled foundation for LLM inference, training, and architecture design: one that is both feasible and beneficial, and a compelling direction for future systems.
title Inference Time Context Sparsity: Illusion or Opportunity?
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.24168