Simulating Hard Attention Using Soft Attention
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866913913183928320 |
|---|---|
| author | Yang, Andy Strobl, Lena Chiang, David Angluin, Dana |
| author_facet | Yang, Andy Strobl, Lena Chiang, David Angluin, Dana |
| contents | We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_09925 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Simulating Hard Attention Using Soft Attention Yang, Andy Strobl, Lena Chiang, David Angluin, Dana Machine Learning Computation and Language Formal Languages and Automata Theory We study conditions under which transformers using soft attention can simulate hard attention, that is, effectively focus all attention on a subset of positions. First, we examine several subclasses of languages recognized by hard-attention transformers, which can be defined in variants of linear temporal logic. We demonstrate how soft-attention transformers can compute formulas of these logics using unbounded positional embeddings or temperature scaling. Second, we demonstrate how temperature scaling allows softmax transformers to simulate general hard-attention transformers, using a temperature that depends on the minimum gap between the maximum attention scores and other attention scores. |
| title | Simulating Hard Attention Using Soft Attention |
| topic | Machine Learning Computation and Language Formal Languages and Automata Theory |
| url | https://arxiv.org/abs/2412.09925 |