Using Span Queries to Optimize for Cache and Attention Locality
Fuente:
arXiv
Saved in:
| Main Authors: | Castro, Paul, Mitchell, Nick, Ordonez, Nathan, Parnell, Thomas, Srivatsa, Mudhakar, Martin, Antoni Viros i |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
by: Joshi, Thomas, et al.
Published: (2025)
by: Joshi, Thomas, et al.
Published: (2025)
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024)
by: Hoque, Adnan, et al.
Published: (2024)
Attack-Resilient Image Watermarking Using Stable Diffusion
by: Zhang, Lijun, et al.
Published: (2024)
by: Zhang, Lijun, et al.
Published: (2024)
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
by: Agarwal, Krish, et al.
Published: (2024)
by: Agarwal, Krish, et al.
Published: (2024)
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025)
by: Devoto, Alessio, et al.
Published: (2025)
SudokuSens: Enhancing Deep Learning Robustness for IoT Sensing Applications using a Generative Approach
by: Wang, Tianshi, et al.
Published: (2024)
by: Wang, Tianshi, et al.
Published: (2024)
Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
by: Li, Allison, et al.
Published: (2025)
by: Li, Allison, et al.
Published: (2025)
Scaling Bidirectional Spans and Span Violations in Attention Mechanism
by: Kim, Jongwook, et al.
Published: (2025)
by: Kim, Jongwook, et al.
Published: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
CacheFormer: High Attention-Based Segment Caching
by: Singh, Sushant, et al.
Published: (2025)
by: Singh, Sushant, et al.
Published: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
by: Yan, Jianxin, et al.
Published: (2026)
by: Yan, Jianxin, et al.
Published: (2026)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
ThinK: Thinner Key Cache by Query-Driven Pruning
by: Xu, Yuhui, et al.
Published: (2024)
by: Xu, Yuhui, et al.
Published: (2024)
Accelerating Production LLMs with Combined Token/Embedding Speculators
by: Wertheimer, Davis, et al.
Published: (2024)
by: Wertheimer, Davis, et al.
Published: (2024)
Multi-Modality Transformer for E-Commerce: Inferring User Purchase Intention to Bridge the Query-Product Gap
by: Mallapragada, Srivatsa, et al.
Published: (2025)
by: Mallapragada, Srivatsa, et al.
Published: (2025)
Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
GPU Performance Portability needs Autotuning
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
Hallucinated Span Detection with Multi-View Attention Features
by: Ogasa, Yuya, et al.
Published: (2025)
by: Ogasa, Yuya, et al.
Published: (2025)
Self-Selected Attention Span for Accelerating Large Language Model Inference
by: Jin, Tian, et al.
Published: (2024)
by: Jin, Tian, et al.
Published: (2024)
From Attribution to Abstention: Training-Free Attention-Based Auditing for Clinical Summarization
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective
by: Li, Anni, et al.
Published: (2025)
by: Li, Anni, et al.
Published: (2025)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
by: Chari, Vivek, et al.
Published: (2025)
by: Chari, Vivek, et al.
Published: (2025)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
by: Fu, Tianyu, et al.
Published: (2024)
by: Fu, Tianyu, et al.
Published: (2024)
Weighted Grouped Query Attention in Transformers
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
by: Chinnakonduru, Sai Sena, et al.
Published: (2024)
Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries
by: Hagar, Nick, et al.
Published: (2025)
by: Hagar, Nick, et al.
Published: (2025)
Randomization Boosts KV Caching, Learning Balances Query Load: A Joint Perspective
by: Wu, Fangzhou, et al.
Published: (2026)
by: Wu, Fangzhou, et al.
Published: (2026)
Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets
by: Xue, Bo, et al.
Published: (2026)
by: Xue, Bo, et al.
Published: (2026)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023)
by: Gim, In, et al.
Published: (2023)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
by: Saxena, Utkarsh, et al.
Published: (2024)
by: Saxena, Utkarsh, et al.
Published: (2024)
An information theoretic approach to quantify the stability of feature selection and ranking algorithms
by: Alaiz-Rodriguez, et al.
Published: (2024)
by: Alaiz-Rodriguez, et al.
Published: (2024)
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories
by: Wang, Jiaming, et al.
Published: (2026)
by: Wang, Jiaming, et al.
Published: (2026)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
Knowledge Localization: Mission Not Accomplished? Enter Query Localization!
by: Chen, Yuheng, et al.
Published: (2024)
by: Chen, Yuheng, et al.
Published: (2024)
The Anatomy of a Triton Attention Kernel
by: Ringlein, Burkhard, et al.
Published: (2025)
by: Ringlein, Burkhard, et al.
Published: (2025)
LoopGuard: Breaking Self-Reinforcing Attention Loops via Dynamic KV Cache Intervention
by: Xu, Dongjie, et al.
Published: (2026)
by: Xu, Dongjie, et al.
Published: (2026)
MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
by: Jiang, Chaoyi, et al.
Published: (2025)
by: Jiang, Chaoyi, et al.
Published: (2025)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
by: Liao, Mengqi, et al.
Published: (2025)
by: Liao, Mengqi, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
Beyond KV Caching: Shared Attention for Efficient LLMs
by: Liao, Bingli, et al.
Published: (2024)
by: Liao, Bingli, et al.
Published: (2024)
Similar Items
-
Paged Attention Meets FlexAttention: Unlocking Long-Context Efficiency in Deployed Inference
by: Joshi, Thomas, et al.
Published: (2025) -
Accelerating a Triton Fused Kernel for W4A16 Quantized Inference with SplitK work decomposition
by: Hoque, Adnan, et al.
Published: (2024) -
Attack-Resilient Image Watermarking Using Stable Diffusion
by: Zhang, Lijun, et al.
Published: (2024) -
HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
by: Agarwal, Krish, et al.
Published: (2024) -
Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
by: Devoto, Alessio, et al.
Published: (2025)