Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xiao, Nachuan, Hu, Xiaoyin, Liu, Xin, Toh, Kim-Chuan
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914684124266496
author Xiao, Nachuan
Hu, Xiaoyin
Liu, Xin
Toh, Kim-Chuan
author_facet Xiao, Nachuan
Hu, Xiaoyin
Liu, Xin
Toh, Kim-Chuan
contents In this paper, we present a comprehensive study on the convergence properties of Adam-family methods for nonsmooth optimization, especially in the training of nonsmooth neural networks. We introduce a novel two-timescale framework that adopts a two-timescale updating scheme, and prove its convergence properties under mild assumptions. Our proposed framework encompasses various popular Adam-family methods, providing convergence guarantees for these methods in training nonsmooth neural networks. Furthermore, we develop stochastic subgradient methods that incorporate gradient clipping techniques for training nonsmooth neural networks with heavy-tailed noise. Through our framework, we show that our proposed methods converge even when the evaluation noises are only assumed to be integrable. Extensive numerical experiments demonstrate the high efficiency and robustness of our proposed methods.
format Preprint
id arxiv_https___arxiv_org_abs_2305_03938
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
Xiao, Nachuan
Hu, Xiaoyin
Liu, Xin
Toh, Kim-Chuan
Optimization and Control
Machine Learning
In this paper, we present a comprehensive study on the convergence properties of Adam-family methods for nonsmooth optimization, especially in the training of nonsmooth neural networks. We introduce a novel two-timescale framework that adopts a two-timescale updating scheme, and prove its convergence properties under mild assumptions. Our proposed framework encompasses various popular Adam-family methods, providing convergence guarantees for these methods in training nonsmooth neural networks. Furthermore, we develop stochastic subgradient methods that incorporate gradient clipping techniques for training nonsmooth neural networks with heavy-tailed noise. Through our framework, we show that our proposed methods converge even when the evaluation noises are only assumed to be integrable. Extensive numerical experiments demonstrate the high efficiency and robustness of our proposed methods.
title Adam-family Methods for Nonsmooth Optimization with Convergence Guarantees
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2305.03938