Perceptrons and localization of attention's mean-field landscape

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Álvarez-López, Antonio, Geshkovski, Borjan, Ruiz-Balet, Domènec
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914562237792256
author Álvarez-López, Antonio
Geshkovski, Borjan
Ruiz-Balet, Domènec
author_facet Álvarez-López, Antonio
Geshkovski, Borjan
Ruiz-Balet, Domènec
contents The forward pass of a Transformer can be seen as an interacting particle system on the unit sphere: time plays the role of layers, particles that of token embeddings, and the unit sphere idealizes layer normalization. In some weight settings the system can even be seen as a gradient flow for an explicit energy, and one can make sense of the infinite context length (mean-field) limit thanks to Wasserstein gradient flows. In this paper we study the effect of the perceptron block in this setting, and show that critical points are generically atomic and localized on subsets of the sphere.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Perceptrons and localization of attention's mean-field landscape
Álvarez-López, Antonio
Geshkovski, Borjan
Ruiz-Balet, Domènec
Machine Learning
Optimization and Control
The forward pass of a Transformer can be seen as an interacting particle system on the unit sphere: time plays the role of layers, particles that of token embeddings, and the unit sphere idealizes layer normalization. In some weight settings the system can even be seen as a gradient flow for an explicit energy, and one can make sense of the infinite context length (mean-field) limit thanks to Wasserstein gradient flows. In this paper we study the effect of the perceptron block in this setting, and show that critical points are generically atomic and localized on subsets of the sphere.
title Perceptrons and localization of attention's mean-field landscape
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2601.21366