Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chezhegov, Savelii, Klyukin, Yaroslav, Semenov, Andrei, Beznosikov, Aleksandr, Gasnikov, Alexander, Horváth, Samuel, Takáč, Martin, Gorbunov, Eduard
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909736110129152
author Chezhegov, Savelii
Klyukin, Yaroslav
Semenov, Andrei
Beznosikov, Aleksandr
Gasnikov, Alexander
Horváth, Samuel
Takáč, Martin
Gorbunov, Eduard
author_facet Chezhegov, Savelii
Klyukin, Yaroslav
Semenov, Andrei
Beznosikov, Aleksandr
Gasnikov, Alexander
Horváth, Samuel
Takáč, Martin
Gorbunov, Eduard
contents Methods with adaptive stepsizes, such as AdaGrad and Adam, are essential for training modern Deep Learning models, especially Large Language Models. Typically, the noise in the stochastic gradients is heavy-tailed for the later ones. Gradient clipping provably helps to achieve good high-probability convergence for such noises. However, despite the similarity between AdaGrad/Adam and Clip-SGD, the current understanding of the high-probability convergence of AdaGrad/Adam-type methods is limited in this case. In this work, we prove that AdaGrad/Adam (and their delayed version) can have provably bad high-probability convergence if the noise is heavy-tailed. We also show that gradient clipping fixes this issue, i.e., we derive new high-probability convergence bounds with polylogarithmic dependence on the confidence level for AdaGrad-Norm and Adam-Norm with clipping and with/without delay for smooth convex/non-convex stochastic optimization with heavy-tailed noise. We extend our results to the case of AdaGrad/Adam with delayed stepsizes. Our empirical evaluations highlight the superiority of clipped versions of AdaGrad/Adam in handling the heavy-tailed noise.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04443
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
Chezhegov, Savelii
Klyukin, Yaroslav
Semenov, Andrei
Beznosikov, Aleksandr
Gasnikov, Alexander
Horváth, Samuel
Takáč, Martin
Gorbunov, Eduard
Machine Learning
Optimization and Control
Methods with adaptive stepsizes, such as AdaGrad and Adam, are essential for training modern Deep Learning models, especially Large Language Models. Typically, the noise in the stochastic gradients is heavy-tailed for the later ones. Gradient clipping provably helps to achieve good high-probability convergence for such noises. However, despite the similarity between AdaGrad/Adam and Clip-SGD, the current understanding of the high-probability convergence of AdaGrad/Adam-type methods is limited in this case. In this work, we prove that AdaGrad/Adam (and their delayed version) can have provably bad high-probability convergence if the noise is heavy-tailed. We also show that gradient clipping fixes this issue, i.e., we derive new high-probability convergence bounds with polylogarithmic dependence on the confidence level for AdaGrad-Norm and Adam-Norm with clipping and with/without delay for smooth convex/non-convex stochastic optimization with heavy-tailed noise. We extend our results to the case of AdaGrad/Adam with delayed stepsizes. Our empirical evaluations highlight the superiority of clipped versions of AdaGrad/Adam in handling the heavy-tailed noise.
title Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2406.04443