Saved in:
Bibliographic Details
Main Authors: Marion, Pierre, Chizat, Lénaïc
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2405.13456
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910671881371648
author Marion, Pierre
Chizat, Lénaïc
author_facet Marion, Pierre
Chizat, Lénaïc
contents The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small one. Indeed, we show a lower bound on the sharpness of minimizers, which grows linearly with depth. We then study the properties of the minimizer found by gradient flow, which is the limit of gradient descent with vanishing learning rate. We show an implicit regularization towards flat minima: the sharpness of the minimizer is no more than a constant times the lower bound. The constant depends on the condition number of the data covariance matrix, but not on width or depth. This result is proven both for a small-scale initialization and a residual initialization. Results of independent interest are shown in both cases. For small-scale initialization, we show that the learned weight matrices are approximately rank-one and that their singular vectors align. For residual initialization, convergence of the gradient flow for a Gaussian initialization of the residual network is proven. Numerical experiments illustrate our results and connect them to gradient descent with non-vanishing learning rate.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13456
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep linear networks for regression are implicitly regularized towards flat minima
Marion, Pierre
Chizat, Lénaïc
Machine Learning
The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics. In this paper, we study the sharpness of deep linear networks for univariate regression. Minimizers can have arbitrarily large sharpness, but not an arbitrarily small one. Indeed, we show a lower bound on the sharpness of minimizers, which grows linearly with depth. We then study the properties of the minimizer found by gradient flow, which is the limit of gradient descent with vanishing learning rate. We show an implicit regularization towards flat minima: the sharpness of the minimizer is no more than a constant times the lower bound. The constant depends on the condition number of the data covariance matrix, but not on width or depth. This result is proven both for a small-scale initialization and a residual initialization. Results of independent interest are shown in both cases. For small-scale initialization, we show that the learned weight matrices are approximately rank-one and that their singular vectors align. For residual initialization, convergence of the gradient flow for a Gaussian initialization of the residual network is proven. Numerical experiments illustrate our results and connect them to gradient descent with non-vanishing learning rate.
title Deep linear networks for regression are implicitly regularized towards flat minima
topic Machine Learning
url https://arxiv.org/abs/2405.13456