Learning Rate Schedules in the Presence of Distribution Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fahrbach, Matthew, Javanmard, Adel, Mirrokni, Vahab, Worah, Pratik
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916289933475840
author Fahrbach, Matthew
Javanmard, Adel
Mirrokni, Vahab
Worah, Pratik
author_facet Fahrbach, Matthew
Javanmard, Adel
Mirrokni, Vahab
Worah, Pratik
contents We design learning rate schedules that minimize regret for SGD-based online learning in the presence of a changing data distribution. We fully characterize the optimal learning rate schedule for online linear regression via a novel analysis with stochastic differential equations. For general convex loss functions, we propose new learning rate schedules that are robust to distribution shift and we give upper and lower bounds for the regret that only differ by constants. For non-convex loss functions, we define a notion of regret based on the gradient norm of the estimated models and propose a learning schedule that minimizes an upper bound on the total expected regret. Intuitively, one expects changing loss landscapes to require more exploration, and we confirm that optimal learning rate schedules typically increase in the presence of distribution shift. Finally, we provide experiments for high-dimensional regression models and neural networks to illustrate these learning rate schedules and their cumulative regret.
format Preprint
id arxiv_https___arxiv_org_abs_2303_15634
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Rate Schedules in the Presence of Distribution Shift
Fahrbach, Matthew
Javanmard, Adel
Mirrokni, Vahab
Worah, Pratik
Machine Learning
Optimization and Control
We design learning rate schedules that minimize regret for SGD-based online learning in the presence of a changing data distribution. We fully characterize the optimal learning rate schedule for online linear regression via a novel analysis with stochastic differential equations. For general convex loss functions, we propose new learning rate schedules that are robust to distribution shift and we give upper and lower bounds for the regret that only differ by constants. For non-convex loss functions, we define a notion of regret based on the gradient norm of the estimated models and propose a learning schedule that minimizes an upper bound on the total expected regret. Intuitively, one expects changing loss landscapes to require more exploration, and we confirm that optimal learning rate schedules typically increase in the presence of distribution shift. Finally, we provide experiments for high-dimensional regression models and neural networks to illustrate these learning rate schedules and their cumulative regret.
title Learning Rate Schedules in the Presence of Distribution Shift
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2303.15634