Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Boursier, Etienne, Pillaud-Vivien, Loucas, Flammarion, Nicolas
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911589428363264
author Boursier, Etienne
Pillaud-Vivien, Loucas
Flammarion, Nicolas
author_facet Boursier, Etienne
Pillaud-Vivien, Loucas
Flammarion, Nicolas
contents The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow dynamics of training one-hidden layer ReLU neural networks for the mean squared error at small initialisation. In this setting, despite non-convexity, we show that the gradient flow converges to zero loss and characterise its implicit bias towards minimum variation norm. Furthermore, some interesting phenomena are highlighted: a quantitative description of the initial alignment phenomenon and a proof that the process follows a specific saddle to saddle dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2206_00939
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
Boursier, Etienne
Pillaud-Vivien, Loucas
Flammarion, Nicolas
Machine Learning
The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow dynamics of training one-hidden layer ReLU neural networks for the mean squared error at small initialisation. In this setting, despite non-convexity, we show that the gradient flow converges to zero loss and characterise its implicit bias towards minimum variation norm. Furthermore, some interesting phenomena are highlighted: a quantitative description of the initial alignment phenomenon and a proof that the process follows a specific saddle to saddle dynamics.
title Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
topic Machine Learning
url https://arxiv.org/abs/2206.00939