Bayes optimal learning of attention-indexed models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Boncoraglio, Fabrizio, Troiani, Emanuele, Erba, Vittorio, Zdeborová, Lenka
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914300008857600
author Boncoraglio, Fabrizio
Troiani, Emanuele
Erba, Vittorio
Zdeborová, Lenka
author_facet Boncoraglio, Fabrizio
Troiani, Emanuele
Erba, Vittorio
Zdeborová, Lenka
contents We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01582
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bayes optimal learning of attention-indexed models
Boncoraglio, Fabrizio
Troiani, Emanuele
Erba, Vittorio
Zdeborová, Lenka
Machine Learning
Disordered Systems and Neural Networks
Information Theory
We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures.
title Bayes optimal learning of attention-indexed models
topic Machine Learning
Disordered Systems and Neural Networks
Information Theory
url https://arxiv.org/abs/2506.01582