On the SDEs and Scaling Rules for Adaptive Gradient Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malladi, Sadhika, Lyu, Kaifeng, Panigrahi, Abhishek, Arora, Sanjeev
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909373961338880
author Malladi, Sadhika
Lyu, Kaifeng
Panigrahi, Abhishek
Arora, Sanjeev
author_facet Malladi, Sadhika
Lyu, Kaifeng
Panigrahi, Abhishek
Arora, Sanjeev
contents Approximating Stochastic Gradient Descent (SGD) as a Stochastic Differential Equation (SDE) has allowed researchers to enjoy the benefits of studying a continuous optimization trajectory while carefully preserving the stochasticity of SGD. Analogous study of adaptive gradient methods, such as RMSprop and Adam, has been challenging because there were no rigorously proven SDE approximations for these methods. This paper derives the SDE approximations for RMSprop and Adam, giving theoretical guarantees of their correctness as well as experimental validation of their applicability to common large-scaling vision and language settings. A key practical result is the derivation of a $\textit{square root scaling rule}$ to adjust the optimization hyperparameters of RMSprop and Adam when changing batch size, and its empirical validation in deep learning settings.
format Preprint
id arxiv_https___arxiv_org_abs_2205_10287
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
Malladi, Sadhika
Lyu, Kaifeng
Panigrahi, Abhishek
Arora, Sanjeev
Machine Learning
Approximating Stochastic Gradient Descent (SGD) as a Stochastic Differential Equation (SDE) has allowed researchers to enjoy the benefits of studying a continuous optimization trajectory while carefully preserving the stochasticity of SGD. Analogous study of adaptive gradient methods, such as RMSprop and Adam, has been challenging because there were no rigorously proven SDE approximations for these methods. This paper derives the SDE approximations for RMSprop and Adam, giving theoretical guarantees of their correctness as well as experimental validation of their applicability to common large-scaling vision and language settings. A key practical result is the derivation of a $\textit{square root scaling rule}$ to adjust the optimization hyperparameters of RMSprop and Adam when changing batch size, and its empirical validation in deep learning settings.
title On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
topic Machine Learning
url https://arxiv.org/abs/2205.10287