Where Do Large Learning Rates Lead Us?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sadrtdinov, Ildus, Kodryan, Maxim, Pokonechny, Eduard, Lobacheva, Ekaterina, Vetrov, Dmitry
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910677245886464
author Sadrtdinov, Ildus
Kodryan, Maxim
Pokonechny, Eduard
Lobacheva, Ekaterina
Vetrov, Dmitry
author_facet Sadrtdinov, Ildus
Kodryan, Maxim
Pokonechny, Eduard
Lobacheva, Ekaterina
Vetrov, Dmitry
contents It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is required for obtaining optimal quality, and 2) what are the key differences between models trained with different LRs? We discover that only a narrow range of initial LRs slightly above the convergence threshold lead to optimal results after fine-tuning with a small LR or weight averaging. By studying the local geometry of reached minima, we observe that using LRs from this optimal range allows for the optimization to locate a basin that only contains high-quality minima. Additionally, we show that these initial LRs result in a sparse set of learned features, with a clear focus on those most relevant for the task. In contrast, starting training with too small LRs leads to unstable minima and attempts to learn all features simultaneously, resulting in poor generalization. Conversely, using initial LRs that are too large fails to detect a basin with good solutions and extract meaningful patterns from the data.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Where Do Large Learning Rates Lead Us?
Sadrtdinov, Ildus
Kodryan, Maxim
Pokonechny, Eduard
Lobacheva, Ekaterina
Vetrov, Dmitry
Machine Learning
It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is required for obtaining optimal quality, and 2) what are the key differences between models trained with different LRs? We discover that only a narrow range of initial LRs slightly above the convergence threshold lead to optimal results after fine-tuning with a small LR or weight averaging. By studying the local geometry of reached minima, we observe that using LRs from this optimal range allows for the optimization to locate a basin that only contains high-quality minima. Additionally, we show that these initial LRs result in a sparse set of learned features, with a clear focus on those most relevant for the task. In contrast, starting training with too small LRs leads to unstable minima and attempts to learn all features simultaneously, resulting in poor generalization. Conversely, using initial LRs that are too large fails to detect a basin with good solutions and extract meaningful patterns from the data.
title Where Do Large Learning Rates Lead Us?
topic Machine Learning
url https://arxiv.org/abs/2410.22113