Adam with model exponential moving average is effective for nonconvex optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahn, Kwangjun, Cutkosky, Ashok
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916460321832960
author Ahn, Kwangjun
Cutkosky, Ashok
author_facet Ahn, Kwangjun
Cutkosky, Ashok
contents In this work, we offer a theoretical analysis of two modern optimization techniques for training large and complex models: (i) adaptive optimization algorithms, such as Adam, and (ii) the model exponential moving average (EMA). Specifically, we demonstrate that a clipped version of Adam with model EMA achieves the optimal convergence rates in various nonconvex optimization settings, both smooth and nonsmooth. Moreover, when the scale varies significantly across different coordinates, we demonstrate that the coordinate-wise adaptivity of Adam is provably advantageous. Notably, unlike previous analyses of Adam, our analysis crucially relies on its core elements -- momentum and discounting factors -- as well as model EMA, motivating their wide applications in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18199
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adam with model exponential moving average is effective for nonconvex optimization
Ahn, Kwangjun
Cutkosky, Ashok
Machine Learning
Optimization and Control
In this work, we offer a theoretical analysis of two modern optimization techniques for training large and complex models: (i) adaptive optimization algorithms, such as Adam, and (ii) the model exponential moving average (EMA). Specifically, we demonstrate that a clipped version of Adam with model EMA achieves the optimal convergence rates in various nonconvex optimization settings, both smooth and nonsmooth. Moreover, when the scale varies significantly across different coordinates, we demonstrate that the coordinate-wise adaptivity of Adam is provably advantageous. Notably, unlike previous analyses of Adam, our analysis crucially relies on its core elements -- momentum and discounting factors -- as well as model EMA, motivating their wide applications in practice.
title Adam with model exponential moving average is effective for nonconvex optimization
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2405.18199