SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Jintao, Xiang, Chendong, Huang, Haofeng, Wei, Jia, Xi, Haocheng, Zhu, Jun, Chen, Jianfei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914163740114944
author Zhang, Jintao
Xiang, Chendong
Huang, Haofeng
Wei, Jia
Xi, Haocheng
Zhu, Jun
Chen, Jianfei
author_facet Zhang, Jintao
Xiang, Chendong
Huang, Haofeng
Wei, Jia
Xi, Haocheng
Zhu, Jun
Chen, Jianfei
contents An efficient attention implementation is essential for large models due to its quadratic time complexity. Fortunately, attention commonly exhibits sparsity, i.e., many values in the attention map are near zero, allowing for the omission of corresponding computations. Many studies have utilized the sparse pattern to accelerate attention. However, most existing works focus on optimizing attention within specific models by exploiting certain sparse patterns of the attention map. A universal sparse attention that guarantees both the speedup and end-to-end performance of diverse models remains elusive. In this paper, we propose SpargeAttn, a universal sparse and quantized attention for any model. Our method uses a two-stage online filter: in the first stage, we rapidly and accurately predict the attention map, enabling the skip of some matrix multiplications in attention. In the second stage, we design an online softmax-aware filter that incurs no extra overhead and further skips some matrix multiplications. Experiments show that our method significantly accelerates diverse models, including language, image, and video generation, without sacrificing end-to-end metrics. The code is available at https://github.com/thu-ml/SpargeAttn.
format Preprint
id arxiv_https___arxiv_org_abs_2502_18137
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
Zhang, Jintao
Xiang, Chendong
Huang, Haofeng
Wei, Jia
Xi, Haocheng
Zhu, Jun
Chen, Jianfei
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Performance
An efficient attention implementation is essential for large models due to its quadratic time complexity. Fortunately, attention commonly exhibits sparsity, i.e., many values in the attention map are near zero, allowing for the omission of corresponding computations. Many studies have utilized the sparse pattern to accelerate attention. However, most existing works focus on optimizing attention within specific models by exploiting certain sparse patterns of the attention map. A universal sparse attention that guarantees both the speedup and end-to-end performance of diverse models remains elusive. In this paper, we propose SpargeAttn, a universal sparse and quantized attention for any model. Our method uses a two-stage online filter: in the first stage, we rapidly and accurately predict the attention map, enabling the skip of some matrix multiplications in attention. In the second stage, we design an online softmax-aware filter that incurs no extra overhead and further skips some matrix multiplications. Experiments show that our method significantly accelerates diverse models, including language, image, and video generation, without sacrificing end-to-end metrics. The code is available at https://github.com/thu-ml/SpargeAttn.
title SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Performance
url https://arxiv.org/abs/2502.18137