Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Titopoulos, Vasileios, Alexandridis, Kosmas, Dimitrakopoulos, Giorgos
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914080472694784
author Titopoulos, Vasileios
Alexandridis, Kosmas
Dimitrakopoulos, Giorgos
author_facet Titopoulos, Vasileios
Alexandridis, Kosmas
Dimitrakopoulos, Giorgos
contents Attention is a core operation in numerous machine learning and artificial intelligence models. This work focuses on the acceleration of attention kernel using FlashAttention algorithm, in vector processors, particularly those based on the RISC-V instruction set architecture (ISA). This work represents the first effort to vectorize FlashAttention, minimizing scalar code and simplifying the computational complexity of evaluating exponentials needed by softmax used in attention. By utilizing a low-cost approximation for exponentials in floating-point arithmetic, we reduce the cost of computing the exponential function without the need to extend baseline vector ISA with new custom instructions. Also, appropriate tiling strategies are explored with the goal to improve memory locality. Experimental results highlight the scalability of our approach, demonstrating significant performance gains with the vectorized implementations when processing attention layers in practical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2510_06834
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
Titopoulos, Vasileios
Alexandridis, Kosmas
Dimitrakopoulos, Giorgos
Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
Attention is a core operation in numerous machine learning and artificial intelligence models. This work focuses on the acceleration of attention kernel using FlashAttention algorithm, in vector processors, particularly those based on the RISC-V instruction set architecture (ISA). This work represents the first effort to vectorize FlashAttention, minimizing scalar code and simplifying the computational complexity of evaluating exponentials needed by softmax used in attention. By utilizing a low-cost approximation for exponentials in floating-point arithmetic, we reduce the cost of computing the exponential function without the need to extend baseline vector ISA with new custom instructions. Also, appropriate tiling strategies are explored with the goal to improve memory locality. Experimental results highlight the scalability of our approach, demonstrating significant performance gains with the vectorized implementations when processing attention layers in practical applications.
title Vectorized FlashAttention with Low-cost Exponential Computation in RISC-V Vector Processors
topic Machine Learning
Distributed, Parallel, and Cluster Computing
Performance
url https://arxiv.org/abs/2510.06834