SystolicAttention: Fusing FlashAttention within a Single Systolic Array

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lin, Jiawei, Li, Yuanlong, Chen, Guokai, Bourgeat, Thomas
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912753770299392
author Lin, Jiawei
Li, Yuanlong
Chen, Guokai
Bourgeat, Thomas
author_facet Lin, Jiawei
Li, Yuanlong
Chen, Guokai
Bourgeat, Thomas
contents Transformer models rely heavily on the scaled dot-product attention (SDPA) operation, typically implemented as FlashAttention. Characterized by its frequent interleaving of matrix multiplications and softmax operations, FlashAttention fails to fully utilize the compute resources of modern systolic-array-based accelerators designed for consecutive and large matrix multiplications. To fully unleash the performance potential of systolic arrays for FlashAttention, we propose FSA, an enhanced systolic array architecture that runs the entire FlashAttention on the array without external vector units. Combined with SystolicAttention, an optimized kernel for FSA that achieves fine-grained and element-wise overlapping of FlashAttention operations, FSA maximizes array utilization while preserving the original floating-point operation order of FlashAttention. We implement FSA in synthesizable RTL and evaluate its performance against state-of-the-art systolic-array-based accelerators. Our results show that FSA achieves 1.77x and 4.83x higher attention FLOPs/s utilization compared to AWS Neuron-v2 and Google TPUv5e, respectively. We synthesize FSA in a 16 nm technology at 1.5 GHz, and results indicate only a 12% area overhead compared to a standard weight-stationary systolic array.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11331
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SystolicAttention: Fusing FlashAttention within a Single Systolic Array
Lin, Jiawei
Li, Yuanlong
Chen, Guokai
Bourgeat, Thomas
Hardware Architecture
Artificial Intelligence
Transformer models rely heavily on the scaled dot-product attention (SDPA) operation, typically implemented as FlashAttention. Characterized by its frequent interleaving of matrix multiplications and softmax operations, FlashAttention fails to fully utilize the compute resources of modern systolic-array-based accelerators designed for consecutive and large matrix multiplications. To fully unleash the performance potential of systolic arrays for FlashAttention, we propose FSA, an enhanced systolic array architecture that runs the entire FlashAttention on the array without external vector units. Combined with SystolicAttention, an optimized kernel for FSA that achieves fine-grained and element-wise overlapping of FlashAttention operations, FSA maximizes array utilization while preserving the original floating-point operation order of FlashAttention. We implement FSA in synthesizable RTL and evaluate its performance against state-of-the-art systolic-array-based accelerators. Our results show that FSA achieves 1.77x and 4.83x higher attention FLOPs/s utilization compared to AWS Neuron-v2 and Google TPUv5e, respectively. We synthesize FSA in a 16 nm technology at 1.5 GHz, and results indicate only a 12% area overhead compared to a standard weight-stationary systolic array.
title SystolicAttention: Fusing FlashAttention within a Single Systolic Array
topic Hardware Architecture
Artificial Intelligence
url https://arxiv.org/abs/2507.11331