Power of Generalized Smoothness in Stochastic Convex Optimization: First- and Zero-Order Algorithms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lobanov, Aleksandr, Gasnikov, Alexander
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912388789305344
author Lobanov, Aleksandr
Gasnikov, Alexander
author_facet Lobanov, Aleksandr
Gasnikov, Alexander
contents This paper is devoted to the study of stochastic optimization problems under the generalized smoothness assumption. By considering the unbiased gradient oracle in Stochastic Gradient Descent, we provide strategies to achieve in bounds the summands describing linear rate. In particular, in the case $L_0 = 0$, we obtain in the convex setup the iteration complexity: $N = \mathcal{O}\left(L_1R \log\frac{1}{\varepsilon} + \frac{L_1 c R^2}{\varepsilon}\right)$ for Clipped Stochastic Gradient Descent and $N = \mathcal{O}\left(L_1R \log\frac{1}{\varepsilon}\right)$ for Normalized Stochastic Gradient Descent. Furthermore, we generalize the convergence results to the case with a biased gradient oracle, and show that the power of $(L_0,L_1)$-smoothness extends to zero-order algorithms. Finally, we demonstrate the possibility of linear convergence in the convex setup through numerical experimentation, which has aroused some interest in the machine learning community.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Power of Generalized Smoothness in Stochastic Convex Optimization: First- and Zero-Order Algorithms
Lobanov, Aleksandr
Gasnikov, Alexander
Optimization and Control
This paper is devoted to the study of stochastic optimization problems under the generalized smoothness assumption. By considering the unbiased gradient oracle in Stochastic Gradient Descent, we provide strategies to achieve in bounds the summands describing linear rate. In particular, in the case $L_0 = 0$, we obtain in the convex setup the iteration complexity: $N = \mathcal{O}\left(L_1R \log\frac{1}{\varepsilon} + \frac{L_1 c R^2}{\varepsilon}\right)$ for Clipped Stochastic Gradient Descent and $N = \mathcal{O}\left(L_1R \log\frac{1}{\varepsilon}\right)$ for Normalized Stochastic Gradient Descent. Furthermore, we generalize the convergence results to the case with a biased gradient oracle, and show that the power of $(L_0,L_1)$-smoothness extends to zero-order algorithms. Finally, we demonstrate the possibility of linear convergence in the convex setup through numerical experimentation, which has aroused some interest in the machine learning community.
title Power of Generalized Smoothness in Stochastic Convex Optimization: First- and Zero-Order Algorithms
topic Optimization and Control
url https://arxiv.org/abs/2501.18198