Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ramapuram, Jason, Danieli, Federico, Dhekane, Eeshan, Weers, Floris, Busbridge, Dan, Ablin, Pierre, Likhomanenko, Tatiana, Digani, Jagrit, Gu, Zijin, Shidani, Amitis, Webb, Russ
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929683768016896
author Ramapuram, Jason
Danieli, Federico
Dhekane, Eeshan
Weers, Floris
Busbridge, Dan
Ablin, Pierre
Likhomanenko, Tatiana
Digani, Jagrit
Gu, Zijin
Shidani, Amitis
Webb, Russ
author_facet Ramapuram, Jason
Danieli, Federico
Dhekane, Eeshan
Weers, Floris
Busbridge, Dan
Ablin, Pierre
Likhomanenko, Tatiana
Digani, Jagrit
Gu, Zijin
Shidani, Amitis
Webb, Russ
contents Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between keys and queries. Recent work has explored alternatives to softmax attention in transformers, such as ReLU and sigmoid activations. In this work, we revisit sigmoid attention and conduct an in-depth theoretical and empirical analysis. Theoretically, we prove that transformers with sigmoid attention are universal function approximators and benefit from improved regularity compared to softmax attention. Through detailed empirical analysis, we identify stabilization of large initial attention norms during the early stages of training as a crucial factor for the successful training of models with sigmoid attention, outperforming prior attempts. We also introduce FLASHSIGMOID, a hardware-aware and memory-efficient implementation of sigmoid attention yielding a 17% inference kernel speed-up over FLASHATTENTION2 on H100 GPUs. Experiments across language, vision, and speech show that properly normalized sigmoid attention matches the strong performance of softmax attention on a wide range of domains and scales, which previous attempts at sigmoid attention were unable to fully achieve. Our work unifies prior art and establishes best practices for sigmoid attention as a drop-in softmax replacement in transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04431
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Ramapuram, Jason
Danieli, Federico
Dhekane, Eeshan
Weers, Floris
Busbridge, Dan
Ablin, Pierre
Likhomanenko, Tatiana
Digani, Jagrit
Gu, Zijin
Shidani, Amitis
Webb, Russ
Machine Learning
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are typically obtained as the softmax of dot products between keys and queries. Recent work has explored alternatives to softmax attention in transformers, such as ReLU and sigmoid activations. In this work, we revisit sigmoid attention and conduct an in-depth theoretical and empirical analysis. Theoretically, we prove that transformers with sigmoid attention are universal function approximators and benefit from improved regularity compared to softmax attention. Through detailed empirical analysis, we identify stabilization of large initial attention norms during the early stages of training as a crucial factor for the successful training of models with sigmoid attention, outperforming prior attempts. We also introduce FLASHSIGMOID, a hardware-aware and memory-efficient implementation of sigmoid attention yielding a 17% inference kernel speed-up over FLASHATTENTION2 on H100 GPUs. Experiments across language, vision, and speech show that properly normalized sigmoid attention matches the strong performance of softmax attention on a wide range of domains and scales, which previous attempts at sigmoid attention were unable to fully achieve. Our work unifies prior art and establishes best practices for sigmoid attention as a drop-in softmax replacement in transformers.
title Theory, Analysis, and Best Practices for Sigmoid Self-Attention
topic Machine Learning
url https://arxiv.org/abs/2409.04431