Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dong, Juechu, Feng, Boyuan, Guessous, Driss, Liang, Yanbo, He, Horace
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916512714981376
author Dong, Juechu
Feng, Boyuan
Guessous, Driss
Liang, Yanbo
He, Horace
author_facet Dong, Juechu
Feng, Boyuan
Guessous, Driss
Liang, Yanbo
He, Horace
contents Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving both the runtime and the memory consumption. However, the importance of FlashAttention combined with its monolithic nature poses a problem for researchers aiming to try new attention variants -- a "software lottery". This problem is exacerbated by the difficulty of writing efficient fused attention kernels, resisting traditional compiler-based approaches. We introduce FlexAttention, a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code. We demonstrate that many existing attention variants (e.g. Alibi, Document Masking, PagedAttention, etc.) can be implemented via FlexAttention, and that we achieve competitive performance compared to these handwritten kernels. Finally, we demonstrate how FlexAttention allows for easy composition of attention variants, solving the combinatorial explosion of attention variants.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05496
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Flex Attention: A Programming Model for Generating Optimized Attention Kernels
Dong, Juechu
Feng, Boyuan
Guessous, Driss
Liang, Yanbo
He, Horace
Machine Learning
Performance
Programming Languages
Over the past 7 years, attention has become one of the most important primitives in deep learning. The primary approach to optimize attention is FlashAttention, which fuses the operation together, drastically improving both the runtime and the memory consumption. However, the importance of FlashAttention combined with its monolithic nature poses a problem for researchers aiming to try new attention variants -- a "software lottery". This problem is exacerbated by the difficulty of writing efficient fused attention kernels, resisting traditional compiler-based approaches. We introduce FlexAttention, a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code. We demonstrate that many existing attention variants (e.g. Alibi, Document Masking, PagedAttention, etc.) can be implemented via FlexAttention, and that we achieve competitive performance compared to these handwritten kernels. Finally, we demonstrate how FlexAttention allows for easy composition of attention variants, solving the combinatorial explosion of attention variants.
title Flex Attention: A Programming Model for Generating Optimized Attention Kernels
topic Machine Learning
Performance
Programming Languages
url https://arxiv.org/abs/2412.05496