Bayes optimal learning of attention-indexed models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914300008857600 |
|---|---|
| author | Boncoraglio, Fabrizio Troiani, Emanuele Erba, Vittorio Zdeborová, Lenka |
| author_facet | Boncoraglio, Fabrizio Troiani, Emanuele Erba, Vittorio Zdeborová, Lenka |
| contents | We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_01582 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Bayes optimal learning of attention-indexed models Boncoraglio, Fabrizio Troiani, Emanuele Erba, Vittorio Zdeborová, Lenka Machine Learning Disordered Systems and Neural Networks Information Theory We introduce the attention-indexed model (AIM), a theoretical framework for analyzing learning in deep attention layers. Inspired by multi-index models, AIM captures how token-level outputs emerge from layered bilinear interactions over high-dimensional embeddings. Unlike prior tractable attention models, AIM allows full-width key and query matrices, aligning more closely with practical transformers. Using tools from statistical mechanics and random matrix theory, we derive closed-form predictions for Bayes-optimal generalization error and identify sharp phase transitions as a function of sample complexity, model width, and sequence length. We propose a matching approximate message passing algorithm and show that gradient descent can reach optimal performance. AIM offers a solvable playground for understanding learning in self-attention layers, that are key components of modern architectures. |
| title | Bayes optimal learning of attention-indexed models |
| topic | Machine Learning Disordered Systems and Neural Networks Information Theory |
| url | https://arxiv.org/abs/2506.01582 |