Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sordello, Matteo, Dalmasso, Niccolò, He, Hangfeng, Su, Weijie
Format: Preprint
Veröffentlicht: 2019
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929246071422976
author Sordello, Matteo
Dalmasso, Niccolò
He, Hangfeng
Su, Weijie
author_facet Sordello, Matteo
Dalmasso, Niccolò
He, Hangfeng
Su, Weijie
contents This paper proposes SplitSGD, a new dynamic learning rate schedule for stochastic optimization. This method decreases the learning rate for better adaptation to the local geometry of the objective function whenever a stationary phase is detected, that is, the iterates are likely to bounce at around a vicinity of a local minimum. The detection is performed by splitting the single thread into two and using the inner product of the gradients from the two threads as a measure of stationarity. Owing to this simple yet provably valid stationarity detection, SplitSGD is easy-to-implement and essentially does not incur additional computational cost than standard SGD. Through a series of extensive experiments, we show that this method is appropriate for both convex problems and training (non-convex) neural networks, with performance compared favorably to other stochastic optimization methods. Importantly, this method is observed to be very robust with a set of default parameters for a wide range of problems and, moreover, can yield better generalization performance than other adaptive gradient methods such as Adam.
format Preprint
id arxiv_https___arxiv_org_abs_1910_08597
institution arXiv
publishDate 2019
record_format arxiv
spellingShingle Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
Sordello, Matteo
Dalmasso, Niccolò
He, Hangfeng
Su, Weijie
Machine Learning
Optimization and Control
Methodology
This paper proposes SplitSGD, a new dynamic learning rate schedule for stochastic optimization. This method decreases the learning rate for better adaptation to the local geometry of the objective function whenever a stationary phase is detected, that is, the iterates are likely to bounce at around a vicinity of a local minimum. The detection is performed by splitting the single thread into two and using the inner product of the gradients from the two threads as a measure of stationarity. Owing to this simple yet provably valid stationarity detection, SplitSGD is easy-to-implement and essentially does not incur additional computational cost than standard SGD. Through a series of extensive experiments, we show that this method is appropriate for both convex problems and training (non-convex) neural networks, with performance compared favorably to other stochastic optimization methods. Importantly, this method is observed to be very robust with a set of default parameters for a wide range of problems and, moreover, can yield better generalization performance than other adaptive gradient methods such as Adam.
title Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
topic Machine Learning
Optimization and Control
Methodology
url https://arxiv.org/abs/1910.08597