Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
Fuente:
arXiv
Saved in:
| Main Authors: | Hassani, Ali, Zhou, Fengzhe, Kane, Aditya, Huang, Jiannan, Chen, Chieh-Yun, Shi, Min, Walton, Steven, Hoehnerbach, Markus, Thakkar, Vijay, Isaev, Michael, Zhang, Qinsheng, Xu, Bing, Wu, Haicheng, Hwu, Wen-mei, Liu, Ming-Yu, Shi, Humphrey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Faster Neighborhood Attention: Reducing the O(n^2) Cost of Self Attention at the Threadblock Level
by: Hassani, Ali, et al.
Published: (2024)
by: Hassani, Ali, et al.
Published: (2024)
Le-DETR: Revisiting Real-Time Detection Transformer with Efficient Encoder Design
by: Huang, Jiannan, et al.
Published: (2026)
by: Huang, Jiannan, et al.
Published: (2026)
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026)
by: Zadouri, Ted, et al.
Published: (2026)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
PAI-Bench: A Comprehensive Benchmark For Physical AI
by: Zhou, Fengzhe, et al.
Published: (2025)
by: Zhou, Fengzhe, et al.
Published: (2025)
AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
by: Jain, Jitesh, et al.
Published: (2025)
by: Jain, Jitesh, et al.
Published: (2025)
VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents
by: Xu, Bing, et al.
Published: (2026)
by: Xu, Bing, et al.
Published: (2026)
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
by: Chen, Chieh-Yun, et al.
Published: (2025)
by: Chen, Chieh-Yun, et al.
Published: (2025)
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
by: Shah, Jay, et al.
Published: (2024)
by: Shah, Jay, et al.
Published: (2024)
Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
by: Lv, Chengtao, et al.
Published: (2026)
by: Lv, Chengtao, et al.
Published: (2026)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
HSR-Enhanced Sparse Attention Acceleration
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
DASSF: Dynamic-Attention Scale-Sequence Fusion for Aerial Object Detection
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
HELIX: Scaling Raw Audio Understanding with Hybrid Mamba-Attention Beyond the Quadratic Limit
by: Khushiyant, et al.
Published: (2026)
by: Khushiyant, et al.
Published: (2026)
On-demand Quick Metasurface Design with Neighborhood Attention Transformer
by: Sun, Zhi, et al.
Published: (2024)
by: Sun, Zhi, et al.
Published: (2024)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
MLHarness: A Scalable Benchmarking System for MLCommons
by: Chang, Yen-Hsiang, et al.
Published: (2021)
by: Chang, Yen-Hsiang, et al.
Published: (2021)
HeartSimSage: Attention-Enhanced Graph Neural Networks for Accelerating Cardiac Mechanics Modeling
by: Shi, Lei, et al.
Published: (2025)
by: Shi, Lei, et al.
Published: (2025)
MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation
by: Lu, Shuaiye, et al.
Published: (2025)
by: Lu, Shuaiye, et al.
Published: (2025)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
by: Park, Jeongmin Brian, et al.
Published: (2023)
by: Park, Jeongmin Brian, et al.
Published: (2023)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
by: Li, Zhonghao, et al.
Published: (2025)
by: Li, Zhonghao, et al.
Published: (2025)
Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
by: Zhao, Mingkuan, et al.
Published: (2025)
by: Zhao, Mingkuan, et al.
Published: (2025)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
by: Jiao, Siyu, et al.
Published: (2024)
by: Jiao, Siyu, et al.
Published: (2024)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
SOL-ExecBench: Speed-of-Light Benchmarking for Real-World GPU Kernels Against Hardware Limits
by: Lin, Edward, et al.
Published: (2026)
by: Lin, Edward, et al.
Published: (2026)
Representation Learning on Heterophilic Graph with Directional Neighborhood Attention
by: Lu, Qincheng, et al.
Published: (2024)
by: Lu, Qincheng, et al.
Published: (2024)
Attention Beyond Neighborhoods: Reviving Transformer for Graph Clustering
by: Xie, Xuanting, et al.
Published: (2025)
by: Xie, Xuanting, et al.
Published: (2025)
$\nabla$NABLA: Neighborhood Adaptive Block-Level Attention
by: Mikhailov, Dmitrii, et al.
Published: (2025)
by: Mikhailov, Dmitrii, et al.
Published: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
by: Qiao, Liang, et al.
Published: (2025)
by: Qiao, Liang, et al.
Published: (2025)
Rectified Sparse Attention
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
AttnMod: Attention-Based New Art Styles
by: Su, Shih-Chieh
Published: (2024)
by: Su, Shih-Chieh
Published: (2024)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification
by: Li, Nian, et al.
Published: (2024)
by: Li, Nian, et al.
Published: (2024)
SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
by: Joshi, Sahil, et al.
Published: (2026)
by: Joshi, Sahil, et al.
Published: (2026)
Graph Triple Attention Network: A Decoupled Perspective
by: Wang, Xiaotang, et al.
Published: (2024)
by: Wang, Xiaotang, et al.
Published: (2024)
FAST: Factorizable Attention for Speeding up Transformers
by: Gerami, Armin, et al.
Published: (2024)
by: Gerami, Armin, et al.
Published: (2024)
Similar Items
-
Faster Neighborhood Attention: Reducing the O(n^2) Cost of Self Attention at the Threadblock Level
by: Hassani, Ali, et al.
Published: (2024) -
Le-DETR: Revisiting Real-Time Detection Transformer with Efficient Encoder Design
by: Huang, Jiannan, et al.
Published: (2026) -
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
by: Zadouri, Ted, et al.
Published: (2026) -
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022) -
PAI-Bench: A Comprehensive Benchmark For Physical AI
by: Zhou, Fengzhe, et al.
Published: (2025)