SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Jintao, Wei, Jia, Zhang, Pengle, Xu, Xiaoming, Huang, Haofeng, Wang, Haoxu, Jiang, Kai, Chen, Jianfei, Zhu, Jun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909990563872768
author Zhang, Jintao
Wei, Jia
Zhang, Pengle
Xu, Xiaoming
Huang, Haofeng
Wang, Haoxu
Jiang, Kai
Chen, Jianfei
Zhu, Jun
author_facet Zhang, Jintao
Wei, Jia
Zhang, Pengle
Xu, Xiaoming
Huang, Haofeng
Wang, Haoxu
Jiang, Kai
Chen, Jianfei
Zhu, Jun
contents The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new FP4 Tensor Cores in Blackwell GPUs to accelerate attention computation. Our implementation achieves 1038 TOPS on RTX5090, which is a 5x speedup over the fastest FlashAttention on RTX5090. Experiments show that our FP4 attention can accelerate inference of various models in a plug-and-play way. Second, we pioneer low-bit attention to training tasks. Existing low-bit attention works like FlashAttention3 and SageAttention focus only on inference. However, the efficiency of training large models is also important. To explore whether low-bit attention can be effectively applied to training tasks, we design an accurate and efficient 8-bit attention for both forward and backward propagation. Experiments indicate that 8-bit attention achieves lossless performance in fine-tuning tasks but exhibits slower convergence in pretraining tasks. The code is available at https://github.com/thu-ml/SageAttention.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11594
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
Zhang, Jintao
Wei, Jia
Zhang, Pengle
Xu, Xiaoming
Huang, Haofeng
Wang, Haoxu
Jiang, Kai
Chen, Jianfei
Zhu, Jun
Machine Learning
Artificial Intelligence
Hardware Architecture
Computer Vision and Pattern Recognition
Performance
The efficiency of attention is important due to its quadratic time complexity. We enhance the efficiency of attention through two key contributions: First, we leverage the new FP4 Tensor Cores in Blackwell GPUs to accelerate attention computation. Our implementation achieves 1038 TOPS on RTX5090, which is a 5x speedup over the fastest FlashAttention on RTX5090. Experiments show that our FP4 attention can accelerate inference of various models in a plug-and-play way. Second, we pioneer low-bit attention to training tasks. Existing low-bit attention works like FlashAttention3 and SageAttention focus only on inference. However, the efficiency of training large models is also important. To explore whether low-bit attention can be effectively applied to training tasks, we design an accurate and efficient 8-bit attention for both forward and backward propagation. Experiments indicate that 8-bit attention achieves lossless performance in fine-tuning tasks but exhibits slower convergence in pretraining tasks. The code is available at https://github.com/thu-ml/SageAttention.
title SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
topic Machine Learning
Artificial Intelligence
Hardware Architecture
Computer Vision and Pattern Recognition
Performance
url https://arxiv.org/abs/2505.11594