Unique Hard Attention: A Tale of Two Sides

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jerad, Selim, Svete, Anej, Li, Jiaoda, Cotterell, Ryan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911262741364736
author Jerad, Selim
Svete, Anej
Li, Jiaoda
Cotterell, Ryan
author_facet Jerad, Selim
Svete, Anej
Li, Jiaoda
Cotterell, Ryan
contents Understanding the expressive power of transformers has recently attracted attention, as it offers insights into their abilities and limitations. Many studies analyze unique hard attention transformers, where attention selects a single position that maximizes the attention scores. When multiple positions achieve the maximum score, either the rightmost or the leftmost of those is chosen. In this paper, we highlight the importance of this seeming triviality. Recently, finite-precision transformers with both leftmost- and rightmost-hard attention were shown to be equivalent to Linear Temporal Logic (LTL). We show that this no longer holds with only leftmost-hard attention -- in that case, they correspond to a \emph{strictly weaker} fragment of LTL. Furthermore, we show that models with leftmost-hard attention are equivalent to \emph{soft} attention, suggesting they may better approximate real-world transformers than right-attention models. These findings refine the landscape of transformer expressivity and underscore the role of attention directionality.
format Preprint
id arxiv_https___arxiv_org_abs_2503_14615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unique Hard Attention: A Tale of Two Sides
Jerad, Selim
Svete, Anej
Li, Jiaoda
Cotterell, Ryan
Machine Learning
Computational Complexity
Computation and Language
Formal Languages and Automata Theory
Understanding the expressive power of transformers has recently attracted attention, as it offers insights into their abilities and limitations. Many studies analyze unique hard attention transformers, where attention selects a single position that maximizes the attention scores. When multiple positions achieve the maximum score, either the rightmost or the leftmost of those is chosen. In this paper, we highlight the importance of this seeming triviality. Recently, finite-precision transformers with both leftmost- and rightmost-hard attention were shown to be equivalent to Linear Temporal Logic (LTL). We show that this no longer holds with only leftmost-hard attention -- in that case, they correspond to a \emph{strictly weaker} fragment of LTL. Furthermore, we show that models with leftmost-hard attention are equivalent to \emph{soft} attention, suggesting they may better approximate real-world transformers than right-attention models. These findings refine the landscape of transformer expressivity and underscore the role of attention directionality.
title Unique Hard Attention: A Tale of Two Sides
topic Machine Learning
Computational Complexity
Computation and Language
Formal Languages and Automata Theory
url https://arxiv.org/abs/2503.14615