Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911589428363264 |
|---|---|
| author | Boursier, Etienne Pillaud-Vivien, Loucas Flammarion, Nicolas |
| author_facet | Boursier, Etienne Pillaud-Vivien, Loucas Flammarion, Nicolas |
| contents | The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow dynamics of training one-hidden layer ReLU neural networks for the mean squared error at small initialisation. In this setting, despite non-convexity, we show that the gradient flow converges to zero loss and characterise its implicit bias towards minimum variation norm. Furthermore, some interesting phenomena are highlighted: a quantitative description of the initial alignment phenomenon and a proof that the process follows a specific saddle to saddle dynamics. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2206_00939 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs Boursier, Etienne Pillaud-Vivien, Loucas Flammarion, Nicolas Machine Learning The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow dynamics of training one-hidden layer ReLU neural networks for the mean squared error at small initialisation. In this setting, despite non-convexity, we show that the gradient flow converges to zero loss and characterise its implicit bias towards minimum variation norm. Furthermore, some interesting phenomena are highlighted: a quantitative description of the initial alignment phenomenon and a proof that the process follows a specific saddle to saddle dynamics. |
| title | Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2206.00939 |