Optimal Linear Decay Learning Rate Schedules and Further Refinements

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Defazio, Aaron, Cutkosky, Ashok, Mehta, Harsh, Mishchenko, Konstantin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913567453741056
author Defazio, Aaron
Cutkosky, Ashok
Mehta, Harsh
Mishchenko, Konstantin
author_facet Defazio, Aaron
Cutkosky, Ashok
Mehta, Harsh
Mishchenko, Konstantin
contents Learning rate schedules used in practice bear little resemblance to those recommended by theory. We close much of this theory/practice gap, and as a consequence are able to derive new problem-adaptive learning rate schedules. Our main technical contribution is a refined analysis of learning rate schedules for a wide class of optimization algorithms (including SGD). When considering only worst-case analysis, our theory predicts that the optimal choice is the linear decay schedule where the step-size is set proportional to 1 - t/T, where t is the current iteration and T is the total number of steps. To go beyond this worst-case analysis, we use the observed gradient norms to derive schedules refined for any particular task. These refined schedules exhibit learning rate warm-up and rapid learning rate annealing near the end of training. Ours is the first systematic approach to automatically yield both of these properties. We perform the most comprehensive evaluation of learning rate schedules to date, evaluating across 10 diverse deep learning problems, a series of LLMs, and a suite of logistic regression problems. We validate that overall, the linear-decay schedule outperforms all commonly used default schedules including cosine annealing. Our adaptive schedule refinement method gives further improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2310_07831
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Optimal Linear Decay Learning Rate Schedules and Further Refinements
Defazio, Aaron
Cutkosky, Ashok
Mehta, Harsh
Mishchenko, Konstantin
Machine Learning
Artificial Intelligence
Learning rate schedules used in practice bear little resemblance to those recommended by theory. We close much of this theory/practice gap, and as a consequence are able to derive new problem-adaptive learning rate schedules. Our main technical contribution is a refined analysis of learning rate schedules for a wide class of optimization algorithms (including SGD). When considering only worst-case analysis, our theory predicts that the optimal choice is the linear decay schedule where the step-size is set proportional to 1 - t/T, where t is the current iteration and T is the total number of steps. To go beyond this worst-case analysis, we use the observed gradient norms to derive schedules refined for any particular task. These refined schedules exhibit learning rate warm-up and rapid learning rate annealing near the end of training. Ours is the first systematic approach to automatically yield both of these properties. We perform the most comprehensive evaluation of learning rate schedules to date, evaluating across 10 diverse deep learning problems, a series of LLMs, and a suite of logistic regression problems. We validate that overall, the linear-decay schedule outperforms all commonly used default schedules including cosine annealing. Our adaptive schedule refinement method gives further improvements.
title Optimal Linear Decay Learning Rate Schedules and Further Refinements
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2310.07831