On the Convergence of Gradient Descent for Large Learning Rates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Crăciun, Alexandru, Ghoshdastidar, Debarghya
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929620207534080
author Crăciun, Alexandru
Ghoshdastidar, Debarghya
author_facet Crăciun, Alexandru
Ghoshdastidar, Debarghya
contents A vast literature on convergence guarantees for gradient descent and derived methods exists at the moment. However, a simple practical situation remains unexplored: when a fixed step size is used, can we expect gradient descent to converge starting from any initialization? We provide fundamental impossibility results showing that convergence becomes impossible no matter the initialization if the step size gets too big. Looking at the asymptotic value of the gradient norm along the optimization trajectory, we see that there is a sharp transition as the step size crosses a critical value. This has been observed by practitioners, yet the true mechanisms through which this happens remain unclear beyond heuristics. Using results from dynamical systems theory, we provide a proof of this in the case of linear neural networks with a squared loss. We also prove the impossibility of convergence for more general losses without requiring strong assumptions such as Lipschitz continuity for the gradient. We validate our findings through experiments with non-linear networks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_13108
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Convergence of Gradient Descent for Large Learning Rates
Crăciun, Alexandru
Ghoshdastidar, Debarghya
Machine Learning
90C26
A vast literature on convergence guarantees for gradient descent and derived methods exists at the moment. However, a simple practical situation remains unexplored: when a fixed step size is used, can we expect gradient descent to converge starting from any initialization? We provide fundamental impossibility results showing that convergence becomes impossible no matter the initialization if the step size gets too big. Looking at the asymptotic value of the gradient norm along the optimization trajectory, we see that there is a sharp transition as the step size crosses a critical value. This has been observed by practitioners, yet the true mechanisms through which this happens remain unclear beyond heuristics. Using results from dynamical systems theory, we provide a proof of this in the case of linear neural networks with a squared loss. We also prove the impossibility of convergence for more general losses without requiring strong assumptions such as Lipschitz continuity for the gradient. We validate our findings through experiments with non-linear networks.
title On the Convergence of Gradient Descent for Large Learning Rates
topic Machine Learning
90C26
url https://arxiv.org/abs/2402.13108