The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Dandi, Yatin, Pesce, Luca, Zdeborová, Lenka, Krzakala, Florent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
by: Dandi, Yatin, et al.
Published: (2024)
by: Dandi, Yatin, et al.
Published: (2024)
Deep Learning of Compositional Targets with Hierarchical Spectral Methods
by: Tabanelli, Hugo, et al.
Published: (2026)
by: Tabanelli, Hugo, et al.
Published: (2026)
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
by: Arnaboldi, Luca, et al.
Published: (2024)
by: Arnaboldi, Luca, et al.
Published: (2024)
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
by: Ghio, Davide, et al.
Published: (2023)
by: Ghio, Davide, et al.
Published: (2023)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
Universality laws for Gaussian mixtures in generalized linear models
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs
by: Arnaboldi, Luca, et al.
Published: (2024)
by: Arnaboldi, Luca, et al.
Published: (2024)
Fundamental computational limits of weak learnability in high-dimensional multi-index models
by: Troiani, Emanuele, et al.
Published: (2024)
by: Troiani, Emanuele, et al.
Published: (2024)
Provable Learning of Random Hierarchy Models and Hierarchical Shallow-to-Deep Chaining
by: Ren, Yunwei, et al.
Published: (2026)
by: Ren, Yunwei, et al.
Published: (2026)
Deep Learning as Neural Low-Degree Filtering: A Spectral Theory of Hierarchical Feature Learning
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Learning with Restricted Boltzmann Machines: Asymptotics of AMP and GD in High Dimensions
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
A Random Matrix Theory Perspective on the Spectrum of Learned Features and Asymptotic Generalization Capabilities
by: Dandi, Yatin, et al.
Published: (2024)
by: Dandi, Yatin, et al.
Published: (2024)
Optimal Spectral Transitions in High-Dimensional Multi-Index Models
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
Rigorous Asymptotics for First-Order Algorithms Through the Dynamical Cavity Method
by: Dandi, Yatin, et al.
Published: (2026)
by: Dandi, Yatin, et al.
Published: (2026)
Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model
by: Tabanelli, Hugo, et al.
Published: (2025)
by: Tabanelli, Hugo, et al.
Published: (2025)
Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
by: Wortsman-Zurich, Arie, et al.
Published: (2026)
Asymptotics of Non-Convex Generalized Linear Models in High-Dimensions: A proof of the replica formula
by: Vilucchio, Matteo, et al.
Published: (2025)
by: Vilucchio, Matteo, et al.
Published: (2025)
Fundamental limits of Non-Linear Low-Rank Matrix Estimation
by: Mergny, Pierre, et al.
Published: (2024)
by: Mergny, Pierre, et al.
Published: (2024)
A phase transition between positional and semantic learning in a solvable model of dot-product attention
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
Analysis of learning a flow-based generative model from limited sample complexity
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
The Nuclear Route: Sharp Asymptotics of ERM in Overparameterized Quadratic Networks
by: Erba, Vittorio, et al.
Published: (2025)
by: Erba, Vittorio, et al.
Published: (2025)
Fundamental Limits of Matrix Sensing: Exact Asymptotics, Universality, and Applications
by: Xu, Yizhou, et al.
Published: (2025)
by: Xu, Yizhou, et al.
Published: (2025)
Maximally-stable Local Optima in Random Graphs and Spin Glasses: Phase Transitions and Universality
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
A Gentle Introduction to Gradient-Based Optimization and Variational Inequalities for Machine Learning
by: Wadia, Neha S., et al.
Published: (2023)
by: Wadia, Neha S., et al.
Published: (2023)
Rigorous dynamical mean field theory for stochastic gradient descent methods
by: Gerbelot, Cedric, et al.
Published: (2022)
by: Gerbelot, Cedric, et al.
Published: (2022)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Gaussian Universality of Perceptrons with Random Labels
by: Gerace, Federica, et al.
Published: (2022)
by: Gerace, Federica, et al.
Published: (2022)
Bayes-optimal learning of an extensive-width neural network from quadratically many samples
by: Maillard, Antoine, et al.
Published: (2024)
by: Maillard, Antoine, et al.
Published: (2024)
The committee machine: Computational to statistical gaps in learning a two-layers neural network
by: Aubin, Benjamin, et al.
Published: (2018)
by: Aubin, Benjamin, et al.
Published: (2018)
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
by: Boncoraglio, Fabrizio, et al.
Published: (2025)
Sequential Dynamics in Ising Spin Glasses
by: Dandi, Yatin, et al.
Published: (2025)
by: Dandi, Yatin, et al.
Published: (2025)
Counting in Small Transformers: The Delicate Interplay between Attention and Feed-Forward Layers
by: Behrens, Freya, et al.
Published: (2024)
by: Behrens, Freya, et al.
Published: (2024)
High-dimensional Asymptotics of Denoising Autoencoders
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
by: Defilippis, Leonardo, et al.
Published: (2025)
by: Defilippis, Leonardo, et al.
Published: (2025)
Dataset distillation for memorized data: Soft labels can leak held-out teacher knowledge
by: Behrens, Freya, et al.
Published: (2025)
by: Behrens, Freya, et al.
Published: (2025)
Learning High-Dimensional Parity Functions with Product Networks using Gradient Descent
by: Larue, Guillaume, et al.
Published: (2026)
by: Larue, Guillaume, et al.
Published: (2026)
A High Dimensional Statistical Model for Adversarial Training: Geometry and Trade-Offs
by: Tanner, Kasimir, et al.
Published: (2024)
by: Tanner, Kasimir, et al.
Published: (2024)
Similar Items
-
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
by: Dandi, Yatin, et al.
Published: (2024) -
Deep Learning of Compositional Targets with Hierarchical Spectral Methods
by: Tabanelli, Hugo, et al.
Published: (2026) -
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
by: Arnaboldi, Luca, et al.
Published: (2024) -
Sampling with flows, diffusion and autoregressive neural networks: A spin-glass perspective
by: Ghio, Davide, et al.
Published: (2023) -
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)