AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yu, Guo, Dong, Wu, Fang, Zhu, Guoliang, Ding, Dian, Zhang, Yiming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
por: Sui, Yueyuan, et al.
Publicado: (2026)
por: Sui, Yueyuan, et al.
Publicado: (2026)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
por: Shan, Jiquan, et al.
Publicado: (2025)
por: Shan, Jiquan, et al.
Publicado: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
por: Gao, Yizhao, et al.
Publicado: (2025)
por: Gao, Yizhao, et al.
Publicado: (2025)
AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers
por: Zhu, Wenhao, et al.
Publicado: (2024)
por: Zhu, Wenhao, et al.
Publicado: (2024)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
por: Chen, Lida, et al.
Publicado: (2025)
por: Chen, Lida, et al.
Publicado: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
vAttention: Verified Sparse Attention
por: Desai, Aditya, et al.
Publicado: (2025)
por: Desai, Aditya, et al.
Publicado: (2025)
RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
por: Xiao, Kaicheng, et al.
Publicado: (2026)
por: Xiao, Kaicheng, et al.
Publicado: (2026)
SALS: Sparse Attention in Latent Space for KV cache Compression
por: Mu, Junlin, et al.
Publicado: (2025)
por: Mu, Junlin, et al.
Publicado: (2025)
Improving Sparse Autoencoder with Dynamic Attention
por: Wang, Dongsheng, et al.
Publicado: (2026)
por: Wang, Dongsheng, et al.
Publicado: (2026)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
por: Li, Zhonghao, et al.
Publicado: (2025)
por: Li, Zhonghao, et al.
Publicado: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
por: Xu, Ceyu, et al.
Publicado: (2026)
por: Xu, Ceyu, et al.
Publicado: (2026)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
por: Yang, Xu, et al.
Publicado: (2026)
por: Yang, Xu, et al.
Publicado: (2026)
Stem: Rethinking Causal Information Flow in Sparse Attention
por: Niu, Lin, et al.
Publicado: (2026)
por: Niu, Lin, et al.
Publicado: (2026)
SchoenbAt: Rethinking Attention with Polynomial basis
por: Guo, Yuhan, et al.
Publicado: (2025)
por: Guo, Yuhan, et al.
Publicado: (2025)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
por: Willette, Jeffrey, et al.
Publicado: (2025)
por: Willette, Jeffrey, et al.
Publicado: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
por: Lee, Heejun, et al.
Publicado: (2023)
por: Lee, Heejun, et al.
Publicado: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
Transformers with Sparse Attention for Granger Causality
por: Mahesh, Riya, et al.
Publicado: (2024)
por: Mahesh, Riya, et al.
Publicado: (2024)
Sparse Attention as Compact Kernel Regression
por: Santos, Saul, et al.
Publicado: (2026)
por: Santos, Saul, et al.
Publicado: (2026)
Trainable Dynamic Mask Sparse Attention
por: Shi, Jingze, et al.
Publicado: (2025)
por: Shi, Jingze, et al.
Publicado: (2025)
ProxyAttn: Guided Sparse Attention via Representative Heads
por: Wang, Yixuan, et al.
Publicado: (2025)
por: Wang, Yixuan, et al.
Publicado: (2025)
Sparse Attention across Multiple-context KV Cache
por: Cao, Ziyi, et al.
Publicado: (2025)
por: Cao, Ziyi, et al.
Publicado: (2025)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
por: Hassani, Ali, et al.
Publicado: (2025)
por: Hassani, Ali, et al.
Publicado: (2025)
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
por: Zhang, Jintao, et al.
Publicado: (2024)
por: Zhang, Jintao, et al.
Publicado: (2024)
Pilot Contamination-Aware Graph Attention Network for Power Control in CFmMIMO
por: Zhang, Tingting, et al.
Publicado: (2025)
por: Zhang, Tingting, et al.
Publicado: (2025)
Diagonal-Tiled Mixed-Precision Attention for Efficient Low-Bit MXFP Inference
por: Ding, Yifu, et al.
Publicado: (2026)
por: Ding, Yifu, et al.
Publicado: (2026)
Scaling Linear Attention with Sparse State Expansion
por: Pan, Yuqi, et al.
Publicado: (2025)
por: Pan, Yuqi, et al.
Publicado: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
por: Piękos, Piotr, et al.
Publicado: (2025)
por: Piękos, Piotr, et al.
Publicado: (2025)
Beyond Classical Attention: Quantum Attention for Scalable Computation
por: Guo, Xuyang, et al.
Publicado: (2023)
por: Guo, Xuyang, et al.
Publicado: (2023)
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
por: Qiu, Quantong, et al.
Publicado: (2026)
por: Qiu, Quantong, et al.
Publicado: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
Multi-Granular Attention based Heterogeneous Hypergraph Neural Network
por: Jin, Hong, et al.
Publicado: (2025)
por: Jin, Hong, et al.
Publicado: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
por: Wen, Kaiyue, et al.
Publicado: (2024)
por: Wen, Kaiyue, et al.
Publicado: (2024)
S2O: Early Stopping for Sparse Attention via Online Permutation
por: Zhang, Yu, et al.
Publicado: (2026)
por: Zhang, Yu, et al.
Publicado: (2026)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
por: Huang, Yuxiang, et al.
Publicado: (2026)
por: Huang, Yuxiang, et al.
Publicado: (2026)
Sparse Masked Attention Policies for Reliable Generalization
por: Horsch, Caroline, et al.
Publicado: (2026)
por: Horsch, Caroline, et al.
Publicado: (2026)
Interpreting Attention Layer Outputs with Sparse Autoencoders
por: Kissane, Connor, et al.
Publicado: (2024)
por: Kissane, Connor, et al.
Publicado: (2024)
Block Sparse Flash Attention
por: Ohayon, Daniel, et al.
Publicado: (2025)
por: Ohayon, Daniel, et al.
Publicado: (2025)
Heterophily-Aware Graph Attention Network
por: Wang, Junfu, et al.
Publicado: (2023)
por: Wang, Junfu, et al.
Publicado: (2023)
Ejemplares similares
-
Modality-Aware Zero-Shot Pruning and Sparse Attention for Efficient Multimodal Edge Inference
por: Sui, Yueyuan, et al.
Publicado: (2026) -
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
por: Shan, Jiquan, et al.
Publicado: (2025) -
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
por: Gao, Yizhao, et al.
Publicado: (2025) -
AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers
por: Zhu, Wenhao, et al.
Publicado: (2024) -
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
por: Chen, Lida, et al.
Publicado: (2025)