DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuxiang, Gonçalves, Nuno M. T., Alvetreti, Federico, Li, Lei, Han, Xu, Ponti, Edoardo M., Martins, André F. T., Treviso, Marcos V. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaSplash-2: Faster Differentiable Sparse Attention
by: Gonçalves, Nuno, et al.
Published: (2026)
by: Gonçalves, Nuno, et al.
Published: (2026)
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025)
by: Gonçalves, Nuno, et al.
Published: (2025)
Sparse Attention as Compact Kernel Regression
by: Santos, Saul, et al.
Published: (2026)
by: Santos, Saul, et al.
Published: (2026)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
EntmaxKV: Support-Aware Decoding for Entmax Attention
by: Duarte, Gonçalo, et al.
Published: (2026)
by: Duarte, Gonçalo, et al.
Published: (2026)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
by: Alvetreti, Federico, et al.
Published: (2025)
by: Alvetreti, Federico, et al.
Published: (2025)
How Effective are State Space Models for Machine Translation?
by: Pitorro, Hugo, et al.
Published: (2024)
by: Pitorro, Hugo, et al.
Published: (2024)
NOSA: Native and Offloadable Sparse Attention
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
by: Pitorro, Hugo, et al.
Published: (2025)
by: Pitorro, Hugo, et al.
Published: (2025)
Sample-efficient Integration of New Modalities into Large Language Models
by: İnce, Osman Batur, et al.
Published: (2025)
by: İnce, Osman Batur, et al.
Published: (2025)
XAttention: Block Sparse Attention with Antidiagonal Scoring
by: Xu, Ruyi, et al.
Published: (2025)
by: Xu, Ruyi, et al.
Published: (2025)
HieraSparse: Hierarchical Semi-Structured Sparse KV Attention
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
Cross-Lingual and Cross-Cultural Variation in Image Descriptions
by: Berger, Uri, et al.
Published: (2024)
by: Berger, Uri, et al.
Published: (2024)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
by: Devoto, Alessio, et al.
Published: (2024)
by: Devoto, Alessio, et al.
Published: (2024)
Scaling Sparse Fine-Tuning to Large Language Models
by: Ansell, Alan, et al.
Published: (2024)
by: Ansell, Alan, et al.
Published: (2024)
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
by: Xu, Yufei, et al.
Published: (2026)
by: Xu, Yufei, et al.
Published: (2026)
Fast Dash
by: Kedar Dabhadkar
Published: (2025)
by: Kedar Dabhadkar
Published: (2025)
Twilight: Adaptive Attention Sparsity with Hierarchical Top-$p$ Pruning
by: Lin, Chaofan, et al.
Published: (2025)
by: Lin, Chaofan, et al.
Published: (2025)
COMPOSIÇÃO DA FLORA ARBÓREA E ARBORESCENTE NO JARDIM BOTÂNICO DE BENTO GONÇALVES, RIO GRANDE DO SUL, BRASIL
by: Bruna Treviso Cenci
Published: (2013)
by: Bruna Treviso Cenci
Published: (2013)
Natively Trainable Sparse Attention for Hierarchical Point Cloud Datasets
by: Lapautre, Nicolas, et al.
Published: (2025)
by: Lapautre, Nicolas, et al.
Published: (2025)
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
by: Leng, Jiaqi, et al.
Published: (2025)
by: Leng, Jiaqi, et al.
Published: (2025)
Hierarchical Attention for Sparse Volumetric Anomaly Detection in Subclinical Keratoconus
by: Kandakji, Lynn, et al.
Published: (2025)
by: Kandakji, Lynn, et al.
Published: (2025)
xTower: A Multilingual LLM for Explaining and Correcting Translation Errors
by: Treviso, Marcos, et al.
Published: (2024)
by: Treviso, Marcos, et al.
Published: (2024)
vAttention: Verified Sparse Attention
by: Desai, Aditya, et al.
Published: (2025)
by: Desai, Aditya, et al.
Published: (2025)
Efficient Sparse Attention needs Adaptive Token Release
by: Zhang, Chaoran, et al.
Published: (2024)
by: Zhang, Chaoran, et al.
Published: (2024)
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
by: Liu, Di, et al.
Published: (2026)
by: Liu, Di, et al.
Published: (2026)
DashCLIP: Leveraging multimodal models for generating semantic embeddings for DoorDash
by: Gurjar, Omkar, et al.
Published: (2025)
by: Gurjar, Omkar, et al.
Published: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Hilbert-Guided Sparse Local Attention
by: Li, Yunge, et al.
Published: (2025)
by: Li, Yunge, et al.
Published: (2025)
HIMOSA: Efficient Remote Sensing Image Super-Resolution with Hierarchical Mixture of Sparse Attention
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
Hierarchical Sparse Attention Framework for Computationally Efficient Classification of Biological Cells
by: Yoshai, Elad, et al.
Published: (2025)
by: Yoshai, Elad, et al.
Published: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
by: Zhao, Weilin, et al.
Published: (2025)
by: Zhao, Weilin, et al.
Published: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
by: Gao, Yizhao, et al.
Published: (2025)
by: Gao, Yizhao, et al.
Published: (2025)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
by: Shen, Zhenyi, et al.
Published: (2025)
by: Shen, Zhenyi, et al.
Published: (2025)
Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
by: Huang, Xingyue, et al.
Published: (2026)
by: Huang, Xingyue, et al.
Published: (2026)
Improving Sparse Autoencoder with Dynamic Attention
by: Wang, Dongsheng, et al.
Published: (2026)
by: Wang, Dongsheng, et al.
Published: (2026)
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
by: Hassani, Ali, et al.
Published: (2025)
by: Hassani, Ali, et al.
Published: (2025)
What's Holding Back Latent Visual Reasoning?
by: Viveiros, André G., et al.
Published: (2026)
by: Viveiros, André G., et al.
Published: (2026)
Similar Items
-
AdaSplash-2: Faster Differentiable Sparse Attention
by: Gonçalves, Nuno, et al.
Published: (2026) -
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025) -
Sparse Attention as Compact Kernel Regression
by: Santos, Saul, et al.
Published: (2026) -
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025) -
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)