Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
Fuente:
arXiv
Saved in:
| Main Authors: | Willette, Jeffrey, Lee, Heejun, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Training-Free Exponential Context Extension via Cascading KV Cache
by: Willette, Jeffrey, et al.
Published: (2024)
by: Willette, Jeffrey, et al.
Published: (2024)
Robust Molecular Property Prediction via Densifying Scarce Labeled Data
by: Kim, Jina, et al.
Published: (2025)
by: Kim, Jina, et al.
Published: (2025)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
Delta Attention Residuals
by: Luo, Cheng, et al.
Published: (2026)
by: Luo, Cheng, et al.
Published: (2026)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
by: Lee, Heejun, et al.
Published: (2025)
by: Lee, Heejun, et al.
Published: (2025)
Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders
by: Steifer, Tomasz
Published: (2026)
by: Steifer, Tomasz
Published: (2026)
MDN: Parallelizing Stepwise Momentum for Delta Linear Attention
by: Huang, Yulong, et al.
Published: (2026)
by: Huang, Yulong, et al.
Published: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
by: Zhang, Jintao, et al.
Published: (2025)
by: Zhang, Jintao, et al.
Published: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
by: Zhang, Jintao, et al.
Published: (2024)
by: Zhang, Jintao, et al.
Published: (2024)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
by: Yang, Lijie, et al.
Published: (2024)
by: Yang, Lijie, et al.
Published: (2024)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
by: Lee, Heejun, et al.
Published: (2024)
by: Lee, Heejun, et al.
Published: (2024)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
Chunk-wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text
by: Xu, Hainan, et al.
Published: (2026)
by: Xu, Hainan, et al.
Published: (2026)
Stochastic Sparse Attention for Memory-Bound Inference
by: Lee, Kyle, et al.
Published: (2026)
by: Lee, Kyle, et al.
Published: (2026)
Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
by: Sobczyk, Aleksandros, et al.
Published: (2026)
by: Sobczyk, Aleksandros, et al.
Published: (2026)
Multi-View Node Pruning for Accurate Graph Representation
by: Kim, Hanjin, et al.
Published: (2025)
by: Kim, Hanjin, et al.
Published: (2025)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
MonarchAttention: Zero-Shot Conversion to Fast, Hardware-Aware Structured Attention
by: Yaras, Can, et al.
Published: (2025)
by: Yaras, Can, et al.
Published: (2025)
Fused3S: Fast Sparse Attention on Tensor Cores
by: Li, Zitong, et al.
Published: (2025)
by: Li, Zitong, et al.
Published: (2025)
MAC-Attention: a Match-Amend-Complete Scheme for Fast and Accurate Attention Computation
by: Yao, Jinghan, et al.
Published: (2026)
by: Yao, Jinghan, et al.
Published: (2026)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
GFlowPO: Generative Flow Network as a Language Model Prompt Optimizer
by: Cho, Junmo, et al.
Published: (2026)
by: Cho, Junmo, et al.
Published: (2026)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
by: Zhou, Ruijie, et al.
Published: (2026)
by: Zhou, Ruijie, et al.
Published: (2026)
DeltaDock: A Unified Framework for Accurate, Efficient, and Physically Reliable Molecular Docking
by: Yan, Jiaxian, et al.
Published: (2024)
by: Yan, Jiaxian, et al.
Published: (2024)
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
by: Sui, Yueyuan, et al.
Published: (2026)
by: Sui, Yueyuan, et al.
Published: (2026)
Gated Delta Networks: Improving Mamba2 with Delta Rule
by: Yang, Songlin, et al.
Published: (2024)
by: Yang, Songlin, et al.
Published: (2024)
Continuous Diffusion Model for Language Modeling
by: Jo, Jaehyeong, et al.
Published: (2025)
by: Jo, Jaehyeong, et al.
Published: (2025)
Generative Modeling on Manifolds Through Mixture of Riemannian Diffusion Processes
by: Jo, Jaehyeong, et al.
Published: (2023)
by: Jo, Jaehyeong, et al.
Published: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)
by: Huang, Xingyue, et al.
Published: (2026)
Fast Inference with Kronecker-Sparse Matrices
by: Gonon, Antoine, et al.
Published: (2024)
by: Gonon, Antoine, et al.
Published: (2024)
Transformers with Sparse Attention for Granger Causality
by: Mahesh, Riya, et al.
Published: (2024)
by: Mahesh, Riya, et al.
Published: (2024)
Sparse Attention as Compact Kernel Regression
by: Santos, Saul, et al.
Published: (2026)
by: Santos, Saul, et al.
Published: (2026)
Drug Discovery with Dynamic Goal-aware Fragments
by: Lee, Seul, et al.
Published: (2023)
by: Lee, Seul, et al.
Published: (2023)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Similar Items
-
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023) -
Training-Free Exponential Context Extension via Cascading KV Cache
by: Willette, Jeffrey, et al.
Published: (2024) -
Robust Molecular Property Prediction via Densifying Scarce Labeled Data
by: Kim, Jina, et al.
Published: (2025) -
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024) -
Delta Attention Residuals
by: Luo, Cheng, et al.
Published: (2026)