Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache
Fuente:
arXiv
Saved in:
| Main Authors: | Dehghankar, Mohsen, Asudeh, Abolfazl |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mining the Minoria: Unknown, Under-represented, and Under-performing Minority Groups
by: Dehghankar, Mohsen, et al.
Published: (2024)
by: Dehghankar, Mohsen, et al.
Published: (2024)
An Efficient Matrix Multiplication Algorithm for Accelerating Inference in Binary and Ternary Neural Networks
by: Dehghankar, Mohsen, et al.
Published: (2024)
by: Dehghankar, Mohsen, et al.
Published: (2024)
Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks
by: Dehghankar, Mohsen, et al.
Published: (2024)
by: Dehghankar, Mohsen, et al.
Published: (2024)
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
by: Dehghankar, Mohsen, et al.
Published: (2026)
by: Dehghankar, Mohsen, et al.
Published: (2026)
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
by: Ebrahimi, Sana, et al.
Published: (2025)
by: Ebrahimi, Sana, et al.
Published: (2025)
HENN: A Hierarchical Epsilon Net Navigation Graph for Approximate Nearest Neighbor Search
by: Dehghankar, Mohsen, et al.
Published: (2025)
by: Dehghankar, Mohsen, et al.
Published: (2025)
Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries
by: Erfanian, Mahdi, et al.
Published: (2024)
by: Erfanian, Mahdi, et al.
Published: (2024)
On Fair Epsilon Net and Geometric Hitting Set
by: Dehghankar, Mohsen, et al.
Published: (2025)
by: Dehghankar, Mohsen, et al.
Published: (2025)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
Fair Set Cover
by: Dehghankar, Mohsen, et al.
Published: (2024)
by: Dehghankar, Mohsen, et al.
Published: (2024)
Dynamic Necklace Splitting
by: Advani, Rishi, et al.
Published: (2025)
by: Advani, Rishi, et al.
Published: (2025)
FlexiCache: Leveraging Temporal Stability of Attention Heads for Efficient KV Cache Management
by: Takbir, Nazmul, et al.
Published: (2025)
by: Takbir, Nazmul, et al.
Published: (2025)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
by: Liu, Guangda, et al.
Published: (2025)
by: Liu, Guangda, et al.
Published: (2025)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks
by: Shen, Hanzhang, et al.
Published: (2026)
by: Shen, Hanzhang, et al.
Published: (2026)
ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
by: Zhang, Qiuyang, et al.
Published: (2026)
by: Zhang, Qiuyang, et al.
Published: (2026)
LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
[Experiments & Analysis] Evaluating the Feasibility of Sampling-Based Techniques for Training Multilayer Perceptrons
by: Ebrahimi, Sana, et al.
Published: (2023)
by: Ebrahimi, Sana, et al.
Published: (2023)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
by: Bai, Yushi, et al.
Published: (2026)
by: Bai, Yushi, et al.
Published: (2026)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
by: Tian, Yuxuan, et al.
Published: (2025)
by: Tian, Yuxuan, et al.
Published: (2025)
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
by: Tang, Hanlin, et al.
Published: (2024)
by: Tang, Hanlin, et al.
Published: (2024)
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
MomentKV: Closing the Directional Gap in KV Cache Eviction for Long-Context Inference
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
by: Yao, Jiayi, et al.
Published: (2026)
by: Yao, Jiayi, et al.
Published: (2026)
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
by: Ye, Lu, et al.
Published: (2024)
by: Ye, Lu, et al.
Published: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference
by: Gokhale, Sai, et al.
Published: (2025)
by: Gokhale, Sai, et al.
Published: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
by: Ramachandran, Akshat, et al.
Published: (2025)
by: Ramachandran, Akshat, et al.
Published: (2025)
Chameleon: Foundation Models for Fairness-aware Multi-modal Data Augmentation to Enhance Coverage of Minorities
by: Erfanian, Mahdi, et al.
Published: (2024)
by: Erfanian, Mahdi, et al.
Published: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
REQUAL-LM: Reliability and Equity through Aggregation in Large Language Models
by: Ebrahimi, Sana, et al.
Published: (2024)
by: Ebrahimi, Sana, et al.
Published: (2024)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference
by: Sun, Hanshi, et al.
Published: (2024)
by: Sun, Hanshi, et al.
Published: (2024)
Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
by: Lesens, Damien, et al.
Published: (2025)
by: Lesens, Damien, et al.
Published: (2025)
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
by: Joo, Donghyeon, et al.
Published: (2025)
by: Joo, Donghyeon, et al.
Published: (2025)
Similar Items
-
Mining the Minoria: Unknown, Under-represented, and Under-performing Minority Groups
by: Dehghankar, Mohsen, et al.
Published: (2024) -
An Efficient Matrix Multiplication Algorithm for Accelerating Inference in Binary and Ternary Neural Networks
by: Dehghankar, Mohsen, et al.
Published: (2024) -
Rank It, Then Ask It: Input Reranking for Maximizing the Performance of LLMs on Symmetric Tasks
by: Dehghankar, Mohsen, et al.
Published: (2024) -
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
by: Dehghankar, Mohsen, et al.
Published: (2026) -
An Adversary-Resistant Multi-Agent LLM System via Credibility Scoring
by: Ebrahimi, Sana, et al.
Published: (2025)