A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Ruinan, Li, Xiao, Yu, Yaoliang, Wang, Baoxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915292926443520
author Jin, Ruinan
Li, Xiao
Yu, Yaoliang
Wang, Baoxiang
author_facet Jin, Ruinan
Li, Xiao
Yu, Yaoliang
Wang, Baoxiang
contents Adaptive Moment Estimation (Adam) is a cornerstone optimization algorithm in deep learning, widely recognized for its flexibility with adaptive learning rates and efficiency in handling large-scale data. However, despite its practical success, the theoretical understanding of Adam's convergence has been constrained by stringent assumptions, such as almost surely bounded stochastic gradients or uniformly bounded gradients, which are more restrictive than those typically required for analyzing stochastic gradient descent (SGD). In this paper, we introduce a novel and comprehensive framework for analyzing the convergence properties of Adam. This framework offers a versatile approach to establishing Adam's convergence. Specifically, we prove that Adam achieves asymptotic (last iterate sense) convergence in both the almost sure sense and the \(L_1\) sense under the relaxed assumptions typically used for SGD, namely \(L\)-smoothness and the ABC inequality. Meanwhile, under the same assumptions, we show that Adam attains non-asymptotic sample complexity bounds similar to those of SGD.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04458
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
Jin, Ruinan
Li, Xiao
Yu, Yaoliang
Wang, Baoxiang
Machine Learning
Optimization and Control
Adaptive Moment Estimation (Adam) is a cornerstone optimization algorithm in deep learning, widely recognized for its flexibility with adaptive learning rates and efficiency in handling large-scale data. However, despite its practical success, the theoretical understanding of Adam's convergence has been constrained by stringent assumptions, such as almost surely bounded stochastic gradients or uniformly bounded gradients, which are more restrictive than those typically required for analyzing stochastic gradient descent (SGD). In this paper, we introduce a novel and comprehensive framework for analyzing the convergence properties of Adam. This framework offers a versatile approach to establishing Adam's convergence. Specifically, we prove that Adam achieves asymptotic (last iterate sense) convergence in both the almost sure sense and the \(L_1\) sense under the relaxed assumptions typically used for SGD, namely \(L\)-smoothness and the ABC inequality. Meanwhile, under the same assumptions, we show that Adam attains non-asymptotic sample complexity bounds similar to those of SGD.
title A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2410.04458