Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Lou, Chao, Jia, Zixia, Zheng, Zilong, Tu, Kewei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sparser, Faster, Lighter Transformer Language Models
by: Cetin, Edoardo, et al.
Published: (2026)
by: Cetin, Edoardo, et al.
Published: (2026)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
by: Zhuang, Xialie, et al.
Published: (2025)
by: Zhuang, Xialie, et al.
Published: (2025)
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
by: Wang, Xiaobo, et al.
Published: (2025)
by: Wang, Xiaobo, et al.
Published: (2025)
AdaSplash-2: Faster Differentiable Sparse Attention
by: Gonçalves, Nuno, et al.
Published: (2026)
by: Gonçalves, Nuno, et al.
Published: (2026)
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
by: Hu, Junhao, et al.
Published: (2026)
by: Hu, Junhao, et al.
Published: (2026)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
by: Chelba, Ciprian, et al.
Published: (2020)
by: Chelba, Ciprian, et al.
Published: (2020)
Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMs
by: Liu, Jinzhe, et al.
Published: (2025)
by: Liu, Jinzhe, et al.
Published: (2025)
Less is More: Parameter-Efficient Selection of Intermediate Tasks for Transfer Learning
by: Schulte, David, et al.
Published: (2024)
by: Schulte, David, et al.
Published: (2024)
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
by: Jo, Dongwon, et al.
Published: (2026)
by: Jo, Dongwon, et al.
Published: (2026)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
The Sparse Frontier: Sparse Attention Trade-offs in Transformer LLMs
by: Nawrot, Piotr, et al.
Published: (2025)
by: Nawrot, Piotr, et al.
Published: (2025)
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
by: Tian, Qiwei, et al.
Published: (2025)
by: Tian, Qiwei, et al.
Published: (2025)
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Dependency Transformer Grammars: Integrating Dependency Structures into Transformer Language Models
by: Zhao, Yida, et al.
Published: (2024)
by: Zhao, Yida, et al.
Published: (2024)
The NLP Task Effectiveness of Long-Range Transformers
by: Qin, Guanghui, et al.
Published: (2022)
by: Qin, Guanghui, et al.
Published: (2022)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
STS: Efficient Sparse Attention with Speculative Token Sparsity
by: Xu, Ceyu, et al.
Published: (2026)
by: Xu, Ceyu, et al.
Published: (2026)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
LIMR: Less is More for RL Scaling
by: Li, Xuefeng, et al.
Published: (2025)
by: Li, Xuefeng, et al.
Published: (2025)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
by: Zhu, Qianchao, et al.
Published: (2024)
by: Zhu, Qianchao, et al.
Published: (2024)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
Hardware-aligned Hierarchical Sparse Attention for Efficient Long-term Memory Access
by: Hu, Xiang, et al.
Published: (2025)
by: Hu, Xiang, et al.
Published: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
by: Wen, Kaiyue, et al.
Published: (2024)
by: Wen, Kaiyue, et al.
Published: (2024)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
by: Wang, Yihan, et al.
Published: (2022)
by: Wang, Yihan, et al.
Published: (2022)
Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
by: Fan, Zehao, et al.
Published: (2025)
by: Fan, Zehao, et al.
Published: (2025)
Less is More: Improving LLM Alignment via Preference Data Selection
by: Deng, Xun, et al.
Published: (2025)
by: Deng, Xun, et al.
Published: (2025)
Star Attention: Efficient LLM Inference over Long Sequences
by: Acharya, Shantanu, et al.
Published: (2024)
by: Acharya, Shantanu, et al.
Published: (2024)
Sparser Block-Sparse Attention via Token Permutation
by: Wang, Xinghao, et al.
Published: (2025)
by: Wang, Xinghao, et al.
Published: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Multipole Attention for Efficient Long Context Reasoning
by: Hooper, Coleman, et al.
Published: (2025)
by: Hooper, Coleman, et al.
Published: (2025)
Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
by: Filipek, Adam
Published: (2025)
by: Filipek, Adam
Published: (2025)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
by: Hagos, Desta Haileselassie, et al.
Published: (2025)
Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts
by: Ye, Jiayuan, et al.
Published: (2026)
by: Ye, Jiayuan, et al.
Published: (2026)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
by: Shrivastava, Vaishnavi, et al.
Published: (2025)
Scaling Linear Attention with Sparse State Expansion
by: Pan, Yuqi, et al.
Published: (2025)
by: Pan, Yuqi, et al.
Published: (2025)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
by: Wu, Bohao, et al.
Published: (2025)
by: Wu, Bohao, et al.
Published: (2025)
Similar Items
-
Sparser, Faster, Lighter Transformer Language Models
by: Cetin, Edoardo, et al.
Published: (2026) -
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
by: Zhuang, Xialie, et al.
Published: (2025) -
Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
by: Wang, Xiaobo, et al.
Published: (2025) -
AdaSplash-2: Faster Differentiable Sparse Attention
by: Gonçalves, Nuno, et al.
Published: (2026) -
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
by: Hu, Junhao, et al.
Published: (2026)