The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dandi, Yatin, Pesce, Luca, Zdeborová, Lenka, Krzakala, Florent
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917079766007808
author Dandi, Yatin
Pesce, Luca
Zdeborová, Lenka
Krzakala, Florent
author_facet Dandi, Yatin
Pesce, Luca
Zdeborová, Lenka
Krzakala, Florent
contents Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of latent subspace dimensionalities. This framework enables us to analytically study the learning dynamics and generalization performance of deep networks compared to shallow ones in the high-dimensional limit. Specifically, our main theorem shows that feature learning with GD successively reduces the effective dimensionality, transforming a high-dimensional problem into a sequence of lower-dimensional ones. This enables learning the target function with drastically less samples than with shallow networks. While the results are proven in a controlled training setting, we also discuss more common training procedures and argue that they learn through the same mechanisms.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13961
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
Dandi, Yatin
Pesce, Luca
Zdeborová, Lenka
Krzakala, Florent
Machine Learning
Understanding the advantages of deep neural networks trained by gradient descent (GD) compared to shallow models remains an open theoretical challenge. In this paper, we introduce a class of target functions (single and multi-index Gaussian hierarchical targets) that incorporate a hierarchy of latent subspace dimensionalities. This framework enables us to analytically study the learning dynamics and generalization performance of deep networks compared to shallow ones in the high-dimensional limit. Specifically, our main theorem shows that feature learning with GD successively reduces the effective dimensionality, transforming a high-dimensional problem into a sequence of lower-dimensional ones. This enables learning the target function with drastically less samples than with shallow networks. While the results are proven in a controlled training setting, we also discuss more common training procedures and argue that they learn through the same mechanisms.
title The Computational Advantage of Depth: Learning High-Dimensional Hierarchical Functions with Gradient Descent
topic Machine Learning
url https://arxiv.org/abs/2502.13961