The emergence of clusters in self-attention dynamics

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Geshkovski, Borjan, Letrouit, Cyril, Polyanskiy, Yury, Rigollet, Philippe
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913232424271872
author Geshkovski, Borjan
Letrouit, Cyril
Polyanskiy, Yury
Rigollet, Philippe
author_facet Geshkovski, Borjan
Letrouit, Cyril
Polyanskiy, Yury
Rigollet, Philippe
contents Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting objects as time tends to infinity. Cluster locations are determined by the initial tokens, confirming context-awareness of representations learned by Transformers. Using techniques from dynamical systems and partial differential equations, we show that the type of limiting object that emerges depends on the spectrum of the value matrix. Additionally, in the one-dimensional case we prove that the self-attention matrix converges to a low-rank Boolean matrix. The combination of these results mathematically confirms the empirical observation made by Vaswani et al. [VSP'17] that leaders appear in a sequence of tokens when processed by Transformers.
format Preprint
id arxiv_https___arxiv_org_abs_2305_05465
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The emergence of clusters in self-attention dynamics
Geshkovski, Borjan
Letrouit, Cyril
Polyanskiy, Yury
Rigollet, Philippe
Machine Learning
Analysis of PDEs
Viewing Transformers as interacting particle systems, we describe the geometry of learned representations when the weights are not time dependent. We show that particles, representing tokens, tend to cluster toward particular limiting objects as time tends to infinity. Cluster locations are determined by the initial tokens, confirming context-awareness of representations learned by Transformers. Using techniques from dynamical systems and partial differential equations, we show that the type of limiting object that emerges depends on the spectrum of the value matrix. Additionally, in the one-dimensional case we prove that the self-attention matrix converges to a low-rank Boolean matrix. The combination of these results mathematically confirms the empirical observation made by Vaswani et al. [VSP'17] that leaders appear in a sequence of tokens when processed by Transformers.
title The emergence of clusters in self-attention dynamics
topic Machine Learning
Analysis of PDEs
url https://arxiv.org/abs/2305.05465