AdaSplash-2: Faster Differentiable Sparse Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Gonçalves, Nuno, Pitorro, Hugo, Niculae, Vlad, Ponti, Edoardo, Li, Lei, Martins, Andre, Treviso, Marcos |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025)
by: Gonçalves, Nuno, et al.
Published: (2025)
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026)
by: Huang, Yuxiang, et al.
Published: (2026)
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025)
by: Vasylenko, Pavlo, et al.
Published: (2025)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
by: Pitorro, Hugo, et al.
Published: (2025)
by: Pitorro, Hugo, et al.
Published: (2025)
Sparse Attention as Compact Kernel Regression
by: Santos, Saul, et al.
Published: (2026)
by: Santos, Saul, et al.
Published: (2026)
How Effective are State Space Models for Machine Translation?
by: Pitorro, Hugo, et al.
Published: (2024)
by: Pitorro, Hugo, et al.
Published: (2024)
The Unreasonable Effectiveness of Random Target Embeddings for Continuous-Output Neural Machine Translation
by: Tokarchuk, Evgeniia, et al.
Published: (2023)
by: Tokarchuk, Evgeniia, et al.
Published: (2023)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
EntmaxKV: Support-Aware Decoding for Entmax Attention
by: Duarte, Gonçalo, et al.
Published: (2026)
by: Duarte, Gonçalo, et al.
Published: (2026)
Angular Dispersion Accelerates $k$-Nearest Neighbors Machine Translation
by: Tokarchuk, Evgeniia, et al.
Published: (2025)
by: Tokarchuk, Evgeniia, et al.
Published: (2025)
Sparse and Structured Hopfield Networks
by: Santos, Saul, et al.
Published: (2024)
by: Santos, Saul, et al.
Published: (2024)
Representation Collapse in Machine Translation Through the Lens of Angular Dispersion
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
by: Tokarchuk, Evgeniia, et al.
Published: (2026)
Scaling Sparse Fine-Tuning to Large Language Models
by: Ansell, Alan, et al.
Published: (2024)
by: Ansell, Alan, et al.
Published: (2024)
Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
by: Lou, Chao, et al.
Published: (2024)
by: Lou, Chao, et al.
Published: (2024)
Hopfield-Fenchel-Young Networks: A Unified Framework for Associative Memory Retrieval
by: Santos, Saul, et al.
Published: (2024)
by: Santos, Saul, et al.
Published: (2024)
Control the Temperature: Selective Sampling for Diverse and High-Quality LLM Outputs
by: Troshin, Sergey, et al.
Published: (2025)
by: Troshin, Sergey, et al.
Published: (2025)
Inference-Time Hyper-Scaling with KV Cache Compression
by: Łańcucki, Adrian, et al.
Published: (2025)
by: Łańcucki, Adrian, et al.
Published: (2025)
Probing the Emergence of Cross-lingual Alignment during LLM Training
by: Wang, Hetong, et al.
Published: (2024)
by: Wang, Hetong, et al.
Published: (2024)
Emergent Communication Pretraining for Few-Shot Machine Translation
by: Li, Yaoyiran, et al.
Published: (2020)
by: Li, Yaoyiran, et al.
Published: (2020)
Sparser, Faster, Lighter Transformer Language Models
by: Cetin, Edoardo, et al.
Published: (2026)
by: Cetin, Edoardo, et al.
Published: (2026)
Keep your distance: learning dispersed embeddings on $\mathbb{S}_m$
by: Tokarchuk, Evgeniia, et al.
Published: (2025)
by: Tokarchuk, Evgeniia, et al.
Published: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
What's Holding Back Latent Visual Reasoning?
by: Viveiros, André G., et al.
Published: (2026)
by: Viveiros, André G., et al.
Published: (2026)
Discrete Latent Structure in Neural Networks
by: Niculae, Vlad, et al.
Published: (2023)
by: Niculae, Vlad, et al.
Published: (2023)
Context-Aware or Context-Insensitive? Assessing LLMs' Performance in Document-Level Translation
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
On Measuring Context Utilization in Document-Level MT Systems
by: Mohammed, Wafaa, et al.
Published: (2024)
by: Mohammed, Wafaa, et al.
Published: (2024)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Mixture of Sparse Attention: Content-Based Learnable Sparse Attention via Expert-Choice Routing
by: Piękos, Piotr, et al.
Published: (2025)
by: Piękos, Piotr, et al.
Published: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
by: Lee, Heejun, et al.
Published: (2023)
by: Lee, Heejun, et al.
Published: (2023)
Mixtures of In-Context Learners
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
by: Chen, Lida, et al.
Published: (2025)
by: Chen, Lida, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
by: Deng, Yichuan, et al.
Published: (2024)
by: Deng, Yichuan, et al.
Published: (2024)
Model Merging by Uncertainty-Based Gradient Matching
by: Daheim, Nico, et al.
Published: (2023)
by: Daheim, Nico, et al.
Published: (2023)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
by: He, Mutian, et al.
Published: (2025)
by: He, Mutian, et al.
Published: (2025)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
Sparse Attention across Multiple-context KV Cache
by: Cao, Ziyi, et al.
Published: (2025)
by: Cao, Ziyi, et al.
Published: (2025)
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
by: Li, Yuhui, et al.
Published: (2024)
by: Li, Yuhui, et al.
Published: (2024)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Similar Items
-
AdaSplash: Adaptive Sparse Flash Attention
by: Gonçalves, Nuno, et al.
Published: (2025) -
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
by: Huang, Yuxiang, et al.
Published: (2026) -
Long-Context Generalization with Sparse Attention
by: Vasylenko, Pavlo, et al.
Published: (2025) -
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
by: Pitorro, Hugo, et al.
Published: (2025) -
Sparse Attention as Compact Kernel Regression
by: Santos, Saul, et al.
Published: (2026)