Inference Time Context Sparsity: Illusion or Opportunity?
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Sahil, Dixit, Prithvi, Chowdhury, Agniva, Shrivastava, Anshumali, Gonzalez, Joseph E., Stoica, Ion, Agrawal, Kumar Krishna, Desai, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
by: Le, Hoang Anh Duy, et al.
Published: (2026)
by: Le, Hoang Anh Duy, et al.
Published: (2026)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
Post-Training Sparse Attention with Double Sparsity
by: Yang, Shuo, et al.
Published: (2024)
by: Yang, Shuo, et al.
Published: (2024)
Speculative Decoding: Performance or Illusion?
by: Liu, Xiaoxuan, et al.
Published: (2025)
by: Liu, Xiaoxuan, et al.
Published: (2025)
IDentity with Locality: An ideal hash for gene sequence search
by: Desai, Aditya, et al.
Published: (2024)
by: Desai, Aditya, et al.
Published: (2024)
K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
by: Cao, Shiyi, et al.
Published: (2026)
by: Cao, Shiyi, et al.
Published: (2026)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Sleep-time Compute: Beyond Inference Scaling at Test-time
by: Lin, Kevin, et al.
Published: (2025)
by: Lin, Kevin, et al.
Published: (2025)
The Time is Here for Just-in-Time Systems: Challenges and Opportunities
by: Liu, Shu, et al.
Published: (2026)
by: Liu, Shu, et al.
Published: (2026)
RedunCut: Measurement-Driven Sampling and Accuracy Performance Modeling for Low-Cost Live Video Analytics
by: Sela, Gur-Eyal, et al.
Published: (2025)
by: Sela, Gur-Eyal, et al.
Published: (2025)
Superintelligent Retrieval Agent: The Next Frontier of Information Retrieval
by: Yang, Zeyu, et al.
Published: (2026)
by: Yang, Zeyu, et al.
Published: (2026)
LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
by: Mishra, Mayank, et al.
Published: (2026)
by: Mishra, Mayank, et al.
Published: (2026)
Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
MPC-Minimized Secure LLM Inference
by: Rathee, Deevashwer, et al.
Published: (2024)
by: Rathee, Deevashwer, et al.
Published: (2024)
A Graph-based Adversarial Imitation Learning Framework for Reliable & Realtime Fleet Scheduling in Urban Air Mobility
by: Poddar, Prithvi, et al.
Published: (2024)
by: Poddar, Prithvi, et al.
Published: (2024)
Empowering Distributed Training with Sparsity-driven Data Synchronization
by: Wang, Zhuang, et al.
Published: (2023)
by: Wang, Zhuang, et al.
Published: (2023)
S*: Test Time Scaling for Code Generation
by: Li, Dacheng, et al.
Published: (2025)
by: Li, Dacheng, et al.
Published: (2025)
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
NEO: Saving GPU Memory Crisis with CPU Offloading for Online LLM Inference
by: Jiang, Xuanlin, et al.
Published: (2024)
by: Jiang, Xuanlin, et al.
Published: (2024)
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
by: Yang, Zeyu, et al.
Published: (2025)
by: Yang, Zeyu, et al.
Published: (2025)
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
by: Cao, Shiyi, et al.
Published: (2024)
by: Cao, Shiyi, et al.
Published: (2024)
KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
Zero Data Retention in LLM-based Enterprise AI Assistants: A Comparative Study of Market Leading Agentic AI Products
by: Gupta, Komal, et al.
Published: (2025)
by: Gupta, Komal, et al.
Published: (2025)
REFRAG: Rethinking RAG based Decoding
by: Lin, Xiaoqiang, et al.
Published: (2025)
by: Lin, Xiaoqiang, et al.
Published: (2025)
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
by: Park, Jongseok, et al.
Published: (2026)
by: Park, Jongseok, et al.
Published: (2026)
MemGPT: Towards LLMs as Operating Systems
by: Packer, Charles, et al.
Published: (2023)
by: Packer, Charles, et al.
Published: (2023)
Why Do Multi-Agent LLM Systems Fail?
by: Cemri, Mert, et al.
Published: (2025)
by: Cemri, Mert, et al.
Published: (2025)
Automated Generation of Diverse Courses of Actions for Multi-Agent Operations using Binary Optimization and Graph Learning
by: Poddar, Prithvi, et al.
Published: (2025)
by: Poddar, Prithvi, et al.
Published: (2025)
Regression-aware Inference with LLMs
by: Lukasik, Michal, et al.
Published: (2024)
by: Lukasik, Michal, et al.
Published: (2024)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
Some Present-Day Problems of Romanian Library Science
by: Stoica, Ion
Published: (1973)
by: Stoica, Ion
Published: (1973)
The Central University Library, Bucharest. Over Seventy-five Years in the History of a Collection
by: Stoica, Ion
Published: (1972)
by: Stoica, Ion
Published: (1972)
Quantum Gatekeeper: Multi-Factor Context-Bound Image Steganography with VQC Based Key Derivation on Quantum Hardware
by: Tomar, Sahil, et al.
Published: (2026)
by: Tomar, Sahil, et al.
Published: (2026)
SageBwd: A Trainable Low-bit Attention
by: Zhang, Jintao, et al.
Published: (2026)
by: Zhang, Jintao, et al.
Published: (2026)
Similar Items
-
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025) -
HashAttention: Semantic Sparsity for Faster Inference
by: Desai, Aditya, et al.
Published: (2024) -
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
by: Joshi, Sahil, et al.
Published: (2026) -
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
by: Le, Hoang Anh Duy, et al.
Published: (2026) -
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)