Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
Fuente:
arXiv
Guardado en:
| Autores principales: | Troiani, Emanuele, Cui, Hugo, Dandi, Yatin, Krzakala, Florent, Zdeborová, Lenka |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fundamental computational limits of weak learnability in high-dimensional multi-index models
por: Troiani, Emanuele, et al.
Publicado: (2024)
por: Troiani, Emanuele, et al.
Publicado: (2024)
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
por: Ghio, Davide, et al.
Publicado: (2023)
por: Ghio, Davide, et al.
Publicado: (2023)
Asymptotics of feature learning in two-layer networks after one gradient-step
por: Cui, Hugo, et al.
Publicado: (2024)
por: Cui, Hugo, et al.
Publicado: (2024)
Bayes-optimal learning of an extensive-width neural network from quadratically many samples
por: Maillard, Antoine, et al.
Publicado: (2024)
por: Maillard, Antoine, et al.
Publicado: (2024)
Bayes optimal learning of attention-indexed models
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
por: Erba, Vittorio, et al.
Publicado: (2025)
por: Erba, Vittorio, et al.
Publicado: (2025)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
por: Dandi, Yatin, et al.
Publicado: (2026)
por: Dandi, Yatin, et al.
Publicado: (2026)
Sequential Dynamics in Ising Spin Glasses
por: Dandi, Yatin, et al.
Publicado: (2025)
por: Dandi, Yatin, et al.
Publicado: (2025)
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
por: Xu, Yizhou, et al.
Publicado: (2025)
por: Xu, Yizhou, et al.
Publicado: (2025)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
por: Dandi, Yatin, et al.
Publicado: (2026)
por: Dandi, Yatin, et al.
Publicado: (2026)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
por: Tabanelli, Hugo, et al.
Publicado: (2025)
por: Tabanelli, Hugo, et al.
Publicado: (2025)
Quenches in the Sherrington-Kirkpatrick model
por: Erba, Vittorio, et al.
Publicado: (2024)
por: Erba, Vittorio, et al.
Publicado: (2024)
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
por: Xu, Yizhou, et al.
Publicado: (2025)
por: Xu, Yizhou, et al.
Publicado: (2025)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
por: Keup, Christian, et al.
Publicado: (2024)
por: Keup, Christian, et al.
Publicado: (2024)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
High-dimensional Asymptotics of Denoising Autoencoders
por: Cui, Hugo, et al.
Publicado: (2023)
por: Cui, Hugo, et al.
Publicado: (2023)
Statistical mechanics of the maximum-average submatrix problem
por: Erba, Vittorio, et al.
Publicado: (2023)
por: Erba, Vittorio, et al.
Publicado: (2023)
On the Atypical Solutions of the Symmetric Binary Perceptron
por: Barbier, Damien, et al.
Publicado: (2023)
por: Barbier, Damien, et al.
Publicado: (2023)
Low-rank Matrix Estimation with Inhomogeneous Noise
por: Guionnet, Alice, et al.
Publicado: (2022)
por: Guionnet, Alice, et al.
Publicado: (2022)
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
por: Erba, Vittorio, et al.
Publicado: (2024)
por: Erba, Vittorio, et al.
Publicado: (2024)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
por: Defilippis, Leonardo, et al.
Publicado: (2025)
por: Defilippis, Leonardo, et al.
Publicado: (2025)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
por: Clarté, Lucas, et al.
Publicado: (2024)
por: Clarté, Lucas, et al.
Publicado: (2024)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
por: Arnaboldi, Luca, et al.
Publicado: (2025)
por: Arnaboldi, Luca, et al.
Publicado: (2025)
The phase diagram of compressed sensing with $\ell_0$-norm regularization
por: Barbier, Damien, et al.
Publicado: (2024)
por: Barbier, Damien, et al.
Publicado: (2024)
The committee machine: Computational to statistical gaps in learning a two-layers neural network
por: Aubin, Benjamin, et al.
Publicado: (2018)
por: Aubin, Benjamin, et al.
Publicado: (2018)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
por: Mendes, Vicente Conde, et al.
Publicado: (2026)
por: Mendes, Vicente Conde, et al.
Publicado: (2026)
Gaussian Universality of Perceptrons with Random Labels
por: Gerace, Federica, et al.
Publicado: (2022)
por: Gerace, Federica, et al.
Publicado: (2022)
Spectral Thresholds in Correlated Spiked Models and Fundamental Limits of Partial Least Squares
por: Mergny, Pierre, et al.
Publicado: (2025)
por: Mergny, Pierre, et al.
Publicado: (2025)
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
por: Dandi, Yatin, et al.
Publicado: (2024)
por: Dandi, Yatin, et al.
Publicado: (2024)
Building Conformal Prediction Intervals with Approximate Message Passing
por: Clarté, Lucas, et al.
Publicado: (2024)
por: Clarté, Lucas, et al.
Publicado: (2024)
Dynamical Cavity Method for Hypergraphs and its Application to Quenches in the k-XOR-SAT Problem
por: Maier, Aude, et al.
Publicado: (2024)
por: Maier, Aude, et al.
Publicado: (2024)
Dynamical Phase Transitions in Graph Cellular Automata
por: Behrens, Freya, et al.
Publicado: (2023)
por: Behrens, Freya, et al.
Publicado: (2023)
High-dimensional learning of narrow neural networks
por: Cui, Hugo
Publicado: (2024)
por: Cui, Hugo
Publicado: (2024)
On the existence of consistent adversarial attacks in high-dimensional linear classification
por: Vilucchio, Matteo, et al.
Publicado: (2025)
por: Vilucchio, Matteo, et al.
Publicado: (2025)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
por: Sagitova, M., et al.
Publicado: (2026)
por: Sagitova, M., et al.
Publicado: (2026)
Minority Takeover in Majority Dynamics: Searching for Rare Initializations via the History Passing Algorithm
por: Jankola, Marek, et al.
Publicado: (2025)
por: Jankola, Marek, et al.
Publicado: (2025)
The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
por: Dandi, Yatin, et al.
Publicado: (2025)
por: Dandi, Yatin, et al.
Publicado: (2025)
Counting and Hardness-of-Finding Fixed Points in Cellular Automata on Random Graphs
por: Koller, Cédric, et al.
Publicado: (2024)
por: Koller, Cédric, et al.
Publicado: (2024)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
por: Farné, Gabriele, et al.
Publicado: (2026)
por: Farné, Gabriele, et al.
Publicado: (2026)
Ejemplares similares
-
Fundamental computational limits of weak learnability in high-dimensional multi-index models
por: Troiani, Emanuele, et al.
Publicado: (2024) -
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
por: Ghio, Davide, et al.
Publicado: (2023) -
Asymptotics of feature learning in two-layer networks after one gradient-step
por: Cui, Hugo, et al.
Publicado: (2024) -
Bayes-optimal learning of an extensive-width neural network from quadratically many samples
por: Maillard, Antoine, et al.
Publicado: (2024) -
Bayes optimal learning of attention-indexed models
por: Boncoraglio, Fabrizio, et al.
Publicado: (2025)