Simulating Hard Attention Using Soft Attention

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Andy, Strobl, Lena, Chiang, David, Angluin, Dana
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913913183928320
author Yang, Andy
Strobl, Lena
Chiang, David
Angluin, Dana
author_facet Yang, Andy
Strobl, Lena
Chiang, David
Angluin, Dana
contents We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Simulating Hard Attention Using Soft Attention
Yang, Andy
Strobl, Lena
Chiang, David
Angluin, Dana
Machine Learning
Computation and Language
Formal Languages and Automata Theory
We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores.
title Simulating Hard Attention Using Soft Attention
topic Machine Learning
Computation and Language
Formal Languages and Automata Theory
url https://arxiv.org/abs/2412.09925