MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sharifnassab, Arsalan, Salehkaleybar, Saber, Sutton, Richard
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916833211187200
author Sharifnassab, Arsalan
Salehkaleybar, Saber
Sutton, Richard
author_facet Sharifnassab, Arsalan
Salehkaleybar, Saber
Sutton, Richard
contents We address the challenge of optimizing meta-parameters (hyperparameters) in machine learning, a key factor for efficient training and high model performance. Rather than relying on expensive meta-parameter search methods, we introduce MetaOptimize: a dynamic approach that adjusts meta-parameters, particularly step sizes (also known as learning rates), during training. More specifically, MetaOptimize can wrap around any first-order optimization algorithm, tuning step sizes on the fly to minimize a specific form of regret that considers the long-term impact of step sizes on training, through a discounted sum of future losses. We also introduce lower-complexity variants of MetaOptimize that, in conjunction with its adaptability to various optimization algorithms, achieve performance comparable to those of the best hand-crafted learning rate schedules across diverse machine learning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02342
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
Sharifnassab, Arsalan
Salehkaleybar, Saber
Sutton, Richard
Machine Learning
Artificial Intelligence
Optimization and Control
We address the challenge of optimizing meta-parameters (hyperparameters) in machine learning, a key factor for efficient training and high model performance. Rather than relying on expensive meta-parameter search methods, we introduce MetaOptimize: a dynamic approach that adjusts meta-parameters, particularly step sizes (also known as learning rates), during training. More specifically, MetaOptimize can wrap around any first-order optimization algorithm, tuning step sizes on the fly to minimize a specific form of regret that considers the long-term impact of step sizes on training, through a discounted sum of future losses. We also introduce lower-complexity variants of MetaOptimize that, in conjunction with its adaptability to various optimization algorithms, achieve performance comparable to those of the best hand-crafted learning rate schedules across diverse machine learning tasks.
title MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters
topic Machine Learning
Artificial Intelligence
Optimization and Control
url https://arxiv.org/abs/2402.02342