Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Duranthon, O., Marion, P., Boyer, C., Loureiro, B., Zdeborová, L. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
by: Duranthon, O., et al.
Published: (2025)
by: Duranthon, O., et al.
Published: (2025)
Asymptotic generalization error of a single-layer graph convolutional network
by: Duranthon, O., et al.
Published: (2024)
by: Duranthon, O., et al.
Published: (2024)
Specialization of softmax attention heads: insights from the high-dimensional single-location model
by: Sagitova, M., et al.
Published: (2026)
by: Sagitova, M., et al.
Published: (2026)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Bilinear Sequence Regression: A Model for Learning from Long Sequences of High-dimensional Tokens
by: Erba, Vittorio, et al.
Published: (2024)
by: Erba, Vittorio, et al.
Published: (2024)
On the existence of consistent adversarial attacks in high-dimensional linear classification
by: Vilucchio, Matteo, et al.
Published: (2025)
by: Vilucchio, Matteo, et al.
Published: (2025)
High-dimensional Asymptotics of Denoising Autoencoders
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Building Conformal Prediction Intervals with Approximate Message Passing
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
by: Keup, Christian, et al.
Published: (2024)
by: Keup, Christian, et al.
Published: (2024)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
by: Farné, Gabriele, et al.
Published: (2026)
by: Farné, Gabriele, et al.
Published: (2026)
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
by: Tanner, Kasimir, et al.
Published: (2024)
by: Tanner, Kasimir, et al.
Published: (2024)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
by: Tabanelli, Hugo, et al.
Published: (2025)
by: Tabanelli, Hugo, et al.
Published: (2025)
Fundamental computational limits of weak learnability in high-dimensional multi-index models
by: Troiani, Emanuele, et al.
Published: (2024)
by: Troiani, Emanuele, et al.
Published: (2024)
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026)
by: Kühn, Marcel, et al.
Published: (2026)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
Spectral Thresholds in Correlated Spiked Models and Fundamental Limits of Partial Least Squares
by: Mergny, Pierre, et al.
Published: (2025)
by: Mergny, Pierre, et al.
Published: (2025)
Gaussian Universality of Perceptrons with Random Labels
by: Gerace, Federica, et al.
Published: (2022)
by: Gerace, Federica, et al.
Published: (2022)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
by: Mendes, Vicente Conde, et al.
Published: (2026)
by: Mendes, Vicente Conde, et al.
Published: (2026)
Dimension-free deterministic equivalents and scaling laws for random feature regression
by: Defilippis, Leonardo, et al.
Published: (2024)
by: Defilippis, Leonardo, et al.
Published: (2024)
Statistical Mechanics of Support Vector Regression
by: Canatar, Abdulkadir, et al.
Published: (2024)
by: Canatar, Abdulkadir, et al.
Published: (2024)
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
by: Huang, Jie, et al.
Published: (2026)
by: Huang, Jie, et al.
Published: (2026)
Bayes optimal learning of attention-indexed models
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
by: Ghio, Davide, et al.
Published: (2023)
by: Ghio, Davide, et al.
Published: (2023)
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
by: Erba, Vittorio, et al.
Published: (2025)
by: Erba, Vittorio, et al.
Published: (2025)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
Learning Linear Regression with Low-Rank Tasks in-Context
by: Takanami, Kaito, et al.
Published: (2025)
by: Takanami, Kaito, et al.
Published: (2025)
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
by: Bordelon, Blake, et al.
Published: (2025)
by: Bordelon, Blake, et al.
Published: (2025)
Statistical Mechanics Calculations Using Variational Autoregressive Networks and Quantum Annealing
by: Tamura, Yuta, et al.
Published: (2024)
by: Tamura, Yuta, et al.
Published: (2024)
Exploring the Energy Landscape of RBMs: Reciprocal Space Insights into Bosons, Hierarchical Learning and Symmetry Breaking
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
by: Toledo-Marin, J. Quetzalcóatl, et al.
Published: (2025)
Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers
by: Tiberi, Lorenzo, et al.
Published: (2024)
by: Tiberi, Lorenzo, et al.
Published: (2024)
Dynamical Mean-Field Theory of Self-Attention Neural Networks
by: Poc-López, Ángel, et al.
Published: (2024)
by: Poc-López, Ángel, et al.
Published: (2024)
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Statistical Physics of Deep Neural Networks: Generalization Capability, Beyond the Infinite Width, and Feature Learning
by: Ariosto, Sebastiano
Published: (2025)
by: Ariosto, Sebastiano
Published: (2025)
Bayes-optimal learning of an extensive-width neural network from quadratically many samples
by: Maillard, Antoine, et al.
Published: (2024)
by: Maillard, Antoine, et al.
Published: (2024)
Similar Items
-
Statistical physics analysis of graph neural networks: Approaching optimality in the contextual stochastic block model
by: Duranthon, O., et al.
Published: (2025) -
Asymptotic generalization error of a single-layer graph convolutional network
by: Duranthon, O., et al.
Published: (2024) -
Specialization of softmax attention heads: insights from the high-dimensional single-location model
by: Sagitova, M., et al.
Published: (2026) -
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025) -
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
by: Clarté, Lucas, et al.
Published: (2024)