Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Troiani, Emanuele, Cui, Hugo, Dandi, Yatin, Krzakala, Florent, Zdeborová, Lenka
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911260060155904
author Troiani, Emanuele
Cui, Hugo
Dandi, Yatin
Krzakala, Florent
Zdeborová, Lenka
author_facet Troiani, Emanuele
Cui, Hugo
Dandi, Yatin
Krzakala, Florent
Zdeborová, Lenka
contents In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayesian-optimal learning, in the limit of large dimension $D$ and commensurably large number of samples $N$, we derive a sharp asymptotic characterization of the optimal performance as well as the performance of the best-known polynomial-time algorithm for this setting --namely approximate message-passing--, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00901
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
Troiani, Emanuele
Cui, Hugo
Dandi, Yatin
Krzakala, Florent
Zdeborová, Lenka
Machine Learning
Disordered Systems and Neural Networks
In this manuscript, we study the learning of deep attention neural networks, defined as the composition of multiple self-attention layers, with tied and low-rank weights. We first establish a mapping of such models to sequence multi-index models, a generalization of the widely studied multi-index model to sequential covariates, for which we establish a number of general results. In the context of Bayesian-optimal learning, in the limit of large dimension $D$ and commensurably large number of samples $N$, we derive a sharp asymptotic characterization of the optimal performance as well as the performance of the best-known polynomial-time algorithm for this setting --namely approximate message-passing--, and characterize sharp thresholds on the minimal sample complexity required for better-than-random prediction performance. Our analysis uncovers, in particular, how the different layers are learned sequentially. Finally, we discuss how this sequential learning can also be observed in a realistic setup.
title Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
topic Machine Learning
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2502.00901