AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yu, Guo, Dong, Wu, Fang, Zhu, Guoliang, Ding, Dian, Zhang, Yiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
von: Sui, Yueyuan, et al.
Veröffentlicht: (2026)
von: Sui, Yueyuan, et al.
Veröffentlicht: (2026)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
von: Shan, Jiquan, et al.
Veröffentlicht: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
von: Gao, Yizhao, et al.
Veröffentlicht: (2025)
von: Gao, Yizhao, et al.
Veröffentlicht: (2025)
AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
von: Chen, Lida, et al.
Veröffentlicht: (2025)
von: Chen, Lida, et al.
Veröffentlicht: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
von: Yuan, Jingyang, et al.
Veröffentlicht: (2025)
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
von: Xiao, Kaicheng, et al.
Veröffentlicht: (2026)
SALS: Sparse Attention in Latent Space for KV cache Compression
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
von: Mu, Junlin, et al.
Veröffentlicht: (2025)
Improving Sparse Autoencoder with Dynamic Attention
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2026)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
von: Li, Zhonghao, et al.
Veröffentlicht: (2025)
von: Li, Zhonghao, et al.
Veröffentlicht: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
von: Xu, Ceyu, et al.
Veröffentlicht: (2026)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
von: Yang, Xu, et al.
Veröffentlicht: (2026)
von: Yang, Xu, et al.
Veröffentlicht: (2026)
Stem: Rethinking Causal Information Flow in Sparse Attention
von: Niu, Lin, et al.
Veröffentlicht: (2026)
von: Niu, Lin, et al.
Veröffentlicht: (2026)
SchoenbAt: Rethinking Attention with Polynomial basis
von: Guo, Yuhan, et al.
Veröffentlicht: (2025)
von: Guo, Yuhan, et al.
Veröffentlicht: (2025)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
von: Willette, Jeffrey, et al.
Veröffentlicht: (2025)
von: Willette, Jeffrey, et al.
Veröffentlicht: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Transformers with Sparse Attention for Granger Causality
von: Mahesh, Riya, et al.
Veröffentlicht: (2024)
von: Mahesh, Riya, et al.
Veröffentlicht: (2024)
Sparse Attention as Compact Kernel Regression
von: Santos, Saul, et al.
Veröffentlicht: (2026)
von: Santos, Saul, et al.
Veröffentlicht: (2026)
Trainable Dynamic Mask Sparse Attention
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
von: Shi, Jingze, et al.
Veröffentlicht: (2025)
ProxyAttn: Guided Sparse Attention via Representative Heads
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Sparse Attention across Multiple-context KV Cache
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
von: Cao, Ziyi, et al.
Veröffentlicht: (2025)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
von: Hassani, Ali, et al.
Veröffentlicht: (2025)
von: Hassani, Ali, et al.
Veröffentlicht: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
von: Zhang, Jintao, et al.
Veröffentlicht: (2024)
Pilot Contamination-Aware Graph Attention Network for Power Control in CFmMIMO
von: Zhang, Tingting, et al.
Veröffentlicht: (2025)
von: Zhang, Tingting, et al.
Veröffentlicht: (2025)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
von: Ding, Yifu, et al.
Veröffentlicht: (2026)
Scaling Linear Attention with Sparse State Expansion
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
von: Pan, Yuqi, et al.
Veröffentlicht: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
Beyond Classical Attention: Quantum Attention for Scalable Computation
von: Guo, Xuyang, et al.
Veröffentlicht: (2023)
von: Guo, Xuyang, et al.
Veröffentlicht: (2023)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
von: Qiu, Quantong, et al.
Veröffentlicht: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
von: Zhang, Jintao, et al.
Veröffentlicht: (2025)
Multi-Granular Attention based Heterogeneous Hypergraph Neural Network
von: Jin, Hong, et al.
Veröffentlicht: (2025)
von: Jin, Hong, et al.
Veröffentlicht: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
S2O: Early Stopping for Sparse Attention via Online Permutation
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yuxiang, et al.
Veröffentlicht: (2026)
Sparse Masked Attention Policies for Reliable Generalization
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
von: Horsch, Caroline, et al.
Veröffentlicht: (2026)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Block Sparse Flash Attention
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
von: Ohayon, Daniel, et al.
Veröffentlicht: (2025)
Heterophily-Aware Graph Attention Network
von: Wang, Junfu, et al.
Veröffentlicht: (2023)
von: Wang, Junfu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
von: Sui, Yueyuan, et al.
Veröffentlicht: (2026) -
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
von: Shan, Jiquan, et al.
Veröffentlicht: (2025) -
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
von: Gao, Yizhao, et al.
Veröffentlicht: (2025) -
AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024) -
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
von: Chen, Lida, et al.
Veröffentlicht: (2025)