Trust Region Methods For Nonconvex Stochastic Optimization Beyond Lipschitz Smoothness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Chenghan, Li, Chenxi, Zhang, Chuwen, Deng, Qi, Ge, Dongdong, Ye, Yinyu
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917914166165504
author Xie, Chenghan
Li, Chenxi
Zhang, Chuwen
Deng, Qi
Ge, Dongdong
Ye, Yinyu
author_facet Xie, Chenghan
Li, Chenxi
Zhang, Chuwen
Deng, Qi
Ge, Dongdong
Ye, Yinyu
contents In many important machine learning applications, the standard assumption of having a globally Lipschitz continuous gradient may fail to hold. This paper delves into a more general $(L_0, L_1)$-smoothness setting, which gains particular significance within the realms of deep neural networks and distributionally robust optimization (DRO). We demonstrate the significant advantage of trust region methods for stochastic nonconvex optimization under such generalized smoothness assumption. We show that first-order trust region methods can recover the normalized and clipped stochastic gradient as special cases and then provide a unified analysis to show their convergence to first-order stationary conditions. Motivated by the important application of DRO, we propose a generalized high-order smoothness condition, under which second-order trust region methods can achieve a complexity of $\mathcal{O}(ε^{-3.5})$ for convergence to second-order stationary points. By incorporating variance reduction, the second-order trust region method obtains an even better complexity of $\mathcal{O}(ε^{-3})$, matching the optimal bound for standard smooth optimization. To our best knowledge, this is the first work to show convergence beyond the first-order stationary condition for generalized smooth optimization. Preliminary experiments show that our proposed algorithms perform favorably compared with existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2310_17319
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Trust Region Methods For Nonconvex Stochastic Optimization Beyond Lipschitz Smoothness
Xie, Chenghan
Li, Chenxi
Zhang, Chuwen
Deng, Qi
Ge, Dongdong
Ye, Yinyu
Optimization and Control
In many important machine learning applications, the standard assumption of having a globally Lipschitz continuous gradient may fail to hold. This paper delves into a more general $(L_0, L_1)$-smoothness setting, which gains particular significance within the realms of deep neural networks and distributionally robust optimization (DRO). We demonstrate the significant advantage of trust region methods for stochastic nonconvex optimization under such generalized smoothness assumption. We show that first-order trust region methods can recover the normalized and clipped stochastic gradient as special cases and then provide a unified analysis to show their convergence to first-order stationary conditions. Motivated by the important application of DRO, we propose a generalized high-order smoothness condition, under which second-order trust region methods can achieve a complexity of $\mathcal{O}(ε^{-3.5})$ for convergence to second-order stationary points. By incorporating variance reduction, the second-order trust region method obtains an even better complexity of $\mathcal{O}(ε^{-3})$, matching the optimal bound for standard smooth optimization. To our best knowledge, this is the first work to show convergence beyond the first-order stationary condition for generalized smooth optimization. Preliminary experiments show that our proposed algorithms perform favorably compared with existing methods.
title Trust Region Methods For Nonconvex Stochastic Optimization Beyond Lipschitz Smoothness
topic Optimization and Control
url https://arxiv.org/abs/2310.17319