Learning time-scales in two-layers neural networks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Berthier, Raphaël, Montanari, Andrea, Zhou, Kangjie
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909548003983360
author Berthier, Raphaël
Montanari, Andrea
Zhou, Kangjie
author_facet Berthier, Raphaël
Montanari, Andrea
Zhou, Kangjie
contents Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averaging over large batches. Long plateaus in which one observes barely any progress alternate with intervals of rapid decrease. These successive phases of learning often take place on very different time scales. Finally, models learnt in an early phase are typically `simpler' or `easier to learn' although in a way that is difficult to formalize. Although theoretical explanations of these phenomena have been put forward, each of them captures at best certain specific regimes. In this paper, we study the gradient flow dynamics of a wide two-layer neural network in high-dimension, when data are distributed according to a single-index model (i.e., the target function depends on a one-dimensional projection of the covariates). Based on a mixture of new rigorous results, non-rigorous mathematical derivations, and numerical simulations, we propose a scenario for the learning dynamics in this setting. In particular, the proposed evolution exhibits separation of timescales and intermittency. These behaviors arise naturally because the population gradient flow can be recast as a singularly perturbed dynamical system.
format Preprint
id arxiv_https___arxiv_org_abs_2303_00055
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning time-scales in two-layers neural networks
Berthier, Raphaël
Montanari, Andrea
Zhou, Kangjie
Machine Learning
Optimization and Control
34E15, 37N40, 68T07
Gradient-based learning in multi-layer neural networks displays a number of striking features. In particular, the decrease rate of empirical risk is non-monotone even after averaging over large batches. Long plateaus in which one observes barely any progress alternate with intervals of rapid decrease. These successive phases of learning often take place on very different time scales. Finally, models learnt in an early phase are typically `simpler' or `easier to learn' although in a way that is difficult to formalize. Although theoretical explanations of these phenomena have been put forward, each of them captures at best certain specific regimes. In this paper, we study the gradient flow dynamics of a wide two-layer neural network in high-dimension, when data are distributed according to a single-index model (i.e., the target function depends on a one-dimensional projection of the covariates). Based on a mixture of new rigorous results, non-rigorous mathematical derivations, and numerical simulations, we propose a scenario for the learning dynamics in this setting. In particular, the proposed evolution exhibits separation of timescales and intermittency. These behaviors arise naturally because the population gradient flow can be recast as a singularly perturbed dynamical system.
title Learning time-scales in two-layers neural networks
topic Machine Learning
Optimization and Control
34E15, 37N40, 68T07
url https://arxiv.org/abs/2303.00055