Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Massucco, Alex, Del Grande, Leonardo, Carioni, Marcello, Brune, Christoph, Schönlieb, Carola-Bibiane
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913146099204096
author Massucco, Alex
Del Grande, Leonardo
Carioni, Marcello
Brune, Christoph
Schönlieb, Carola-Bibiane
author_facet Massucco, Alex
Del Grande, Leonardo
Carioni, Marcello
Brune, Christoph
Schönlieb, Carola-Bibiane
contents In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretical point of view, the mathematical models proposed in the literature often lack direct contact with the actual architectures and depend on strong simplifying assumptions. In this paper, we reduce this gap by modelling the data flow in multi-headed transformer architectures as time-dependent gradient flows for a suitable interaction energy capturing the design of the attention mechanism. The explicit dependence on time allows us to consider different weights for each head and for each layer, without imposing constraints on the initialization method. Moreover, we prove that, under a suitable integrability assumption on the evolution of the weights, each element of the $ω$-limit set of the gradient flows is a stationary point of the interaction energy at a limiting weight distribution. Finally, we analyse the stability of the gradient flows considering perturbations of both the initial data and the weights. Specifically, on the one hand, we study the robustness of the proposed models with respect to noisy inputs, establishing a continuous dependence of the gradient flows on the initial data and uniqueness of the flows. On the other hand, we prove the $Γ$-convergence of the perturbed interaction energy to the unperturbed one, leading to the convergence of the corresponding gradient flows. We complement these theoretical results with numerical experiments that confirm the predicted energy-dissipation identity and clarify the asymptotic behavior of the dynamics in both the autonomous-like (Ornstein--Uhlenbeck) and the genuinely non-autonomous (oscillating-weights) regimes.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18870
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
Massucco, Alex
Del Grande, Leonardo
Carioni, Marcello
Brune, Christoph
Schönlieb, Carola-Bibiane
Machine Learning
Analysis of PDEs
Functional Analysis
68T07, 35Q93, 49J53, 46N10, 49Q20
In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretical point of view, the mathematical models proposed in the literature often lack direct contact with the actual architectures and depend on strong simplifying assumptions. In this paper, we reduce this gap by modelling the data flow in multi-headed transformer architectures as time-dependent gradient flows for a suitable interaction energy capturing the design of the attention mechanism. The explicit dependence on time allows us to consider different weights for each head and for each layer, without imposing constraints on the initialization method. Moreover, we prove that, under a suitable integrability assumption on the evolution of the weights, each element of the $ω$-limit set of the gradient flows is a stationary point of the interaction energy at a limiting weight distribution. Finally, we analyse the stability of the gradient flows considering perturbations of both the initial data and the weights. Specifically, on the one hand, we study the robustness of the proposed models with respect to noisy inputs, establishing a continuous dependence of the gradient flows on the initial data and uniqueness of the flows. On the other hand, we prove the $Γ$-convergence of the perturbed interaction energy to the unperturbed one, leading to the convergence of the corresponding gradient flows. We complement these theoretical results with numerical experiments that confirm the predicted energy-dissipation identity and clarify the asymptotic behavior of the dynamics in both the autonomous-like (Ornstein--Uhlenbeck) and the genuinely non-autonomous (oscillating-weights) regimes.
title Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows
topic Machine Learning
Analysis of PDEs
Functional Analysis
68T07, 35Q93, 49J53, 46N10, 49Q20
url https://arxiv.org/abs/2605.18870