Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barboni, Raphaël, Peyré, Gabriel, Vialard, François-Xavier
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911067342372864
author Barboni, Raphaël
Peyré, Gabriel
Vialard, François-Xavier
author_facet Barboni, Raphaël
Peyré, Gabriel
Vialard, François-Xavier
contents We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimization problem, most known convergence results are either qualitative or rely on a neural tangent kernel analysis where nonlinear representations of the data are fixed. Using that this problem belongs to the class of separable nonlinear least squares problems, we consider here a Variable Projection (VarPro) or two-timescale learning algorithm, thereby eliminating the linear variables and reducing the learning problem to the training of nonlinear features. In a teacher-student scenario, we show such a strategy enables provable convergence rates for the sampling of a teacher feature distribution. Precisely, in the limit where the regularization strength vanishes, we show that the dynamic of the feature distribution corresponds to a weighted ultra-fast diffusion equation. Recent results on the asymptotic behavior of such PDEs then give quantitative guarantees for the convergence of the learned feature distribution.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18208
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
Barboni, Raphaël
Peyré, Gabriel
Vialard, François-Xavier
Machine Learning
Optimization and Control
We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimization problem, most known convergence results are either qualitative or rely on a neural tangent kernel analysis where nonlinear representations of the data are fixed. Using that this problem belongs to the class of separable nonlinear least squares problems, we consider here a Variable Projection (VarPro) or two-timescale learning algorithm, thereby eliminating the linear variables and reducing the learning problem to the training of nonlinear features. In a teacher-student scenario, we show such a strategy enables provable convergence rates for the sampling of a teacher feature distribution. Precisely, in the limit where the regularization strength vanishes, we show that the dynamic of the feature distribution corresponds to a weighted ultra-fast diffusion equation. Recent results on the asymptotic behavior of such PDEs then give quantitative guarantees for the convergence of the learned feature distribution.
title Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2504.18208