The Road Less Scheduled

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Defazio, Aaron, Yang, Xingyu Alice, Mehta, Harsh, Mishchenko, Konstantin, Khaled, Ahmed, Cutkosky, Ashok
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910678072164352
author Defazio, Aaron
Yang, Xingyu Alice
Mehta, Harsh
Mishchenko, Konstantin
Khaled, Ahmed
Cutkosky, Ashok
author_facet Defazio, Aaron
Yang, Xingyu Alice
Mehta, Harsh
Mishchenko, Konstantin
Khaled, Ahmed
Cutkosky, Ashok
contents Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15682
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Road Less Scheduled
Defazio, Aaron
Yang, Xingyu Alice
Mehta, Harsh
Mishchenko, Konstantin
Khaled, Ahmed
Cutkosky, Ashok
Machine Learning
Artificial Intelligence
Optimization and Control
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We propose an approach that avoids the need for this stopping time by eschewing the use of schedules entirely, while exhibiting state-of-the-art performance compared to schedules across a wide family of problems ranging from convex problems to large-scale deep learning problems. Our Schedule-Free approach introduces no additional hyper-parameters over standard optimizers with momentum. Our method is a direct consequence of a new theory we develop that unifies scheduling and iterate averaging. An open source implementation of our method is available at https://github.com/facebookresearch/schedule_free. Schedule-Free AdamW is the core algorithm behind our winning entry to the MLCommons 2024 AlgoPerf Algorithmic Efficiency Challenge Self-Tuning track.
title The Road Less Scheduled
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2405.15682