Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Burger, Martin, Kabri, Samira, Korolev, Yury, Roith, Tim, Weigand, Lukas
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912348540764160
author Burger, Martin
Kabri, Samira
Korolev, Yury
Roith, Tim
Weigand, Lukas
author_facet Burger, Martin
Kabri, Samira
Korolev, Yury
Roith, Tim
Weigand, Lukas
contents The aim of this paper is to provide a mathematical analysis of transformer architectures using a self-attention mechanism with layer normalization. In particular, observed patterns in such architectures resembling either clusters or uniform distributions pose a number of challenging mathematical questions. We focus on a special case that admits a gradient flow formulation in the spaces of probability measures on the unit sphere under a special metric, which allows us to give at least partial answers in a rigorous way. The arising mathematical problems resemble those recently studied in aggregation equations, but with additional challenges emerging from restricting the dynamics to the sphere and the particular form of the interaction energy. We provide a rigorous framework for studying the gradient flow, which also suggests a possible metric geometry to study the general case (i.e. one that is not described by a gradient flow). We further analyze the stationary points of the induced self-attention dynamics. The latter are related to stationary points of the interaction energy in the Wasserstein geometry, and we further discuss energy minimizers and maximizers in different parameter settings.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03096
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization
Burger, Martin
Kabri, Samira
Korolev, Yury
Roith, Tim
Weigand, Lukas
Analysis of PDEs
68Q32, 49Q20, 43A35
The aim of this paper is to provide a mathematical analysis of transformer architectures using a self-attention mechanism with layer normalization. In particular, observed patterns in such architectures resembling either clusters or uniform distributions pose a number of challenging mathematical questions. We focus on a special case that admits a gradient flow formulation in the spaces of probability measures on the unit sphere under a special metric, which allows us to give at least partial answers in a rigorous way. The arising mathematical problems resemble those recently studied in aggregation equations, but with additional challenges emerging from restricting the dynamics to the sphere and the particular form of the interaction energy. We provide a rigorous framework for studying the gradient flow, which also suggests a possible metric geometry to study the general case (i.e. one that is not described by a gradient flow). We further analyze the stationary points of the induced self-attention dynamics. The latter are related to stationary points of the interaction energy in the Wasserstein geometry, and we further discuss energy minimizers and maximizers in different parameter settings.
title Analysis of mean-field models arising from self-attention dynamics in transformer architectures with layer normalization
topic Analysis of PDEs
68Q32, 49Q20, 43A35
url https://arxiv.org/abs/2501.03096