Learning-rate-free Momentum SGD with Reshuffling Converges in Nonsmooth Nonconvex Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Xiaoyin, Xiao, Nachuan, Liu, Xin, Toh, Kim-Chuan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910502675808256
author Hu, Xiaoyin
Xiao, Nachuan
Liu, Xin
Toh, Kim-Chuan
author_facet Hu, Xiaoyin
Xiao, Nachuan
Liu, Xin
Toh, Kim-Chuan
contents In this paper, we propose a generalized framework for developing learning-rate-free momentum stochastic gradient descent (SGD) methods in the minimization of nonsmooth nonconvex functions, especially in training nonsmooth neural networks. Our framework adaptively generates learning rates based on the historical data of stochastic subgradients and iterates. Under mild conditions, we prove that our proposed framework enjoys global convergence to the stationary points of the objective function in the sense of the conservative field, hence providing convergence guarantees for training nonsmooth neural networks. Based on our proposed framework, we propose a novel learning-rate-free momentum SGD method (LFM). Preliminary numerical experiments reveal that LFM performs comparably to the state-of-the-art learning-rate-free methods (which have not been shown theoretically to be convergence) across well-known neural network training benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18287
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning-rate-free Momentum SGD with Reshuffling Converges in Nonsmooth Nonconvex Optimization
Hu, Xiaoyin
Xiao, Nachuan
Liu, Xin
Toh, Kim-Chuan
Optimization and Control
In this paper, we propose a generalized framework for developing learning-rate-free momentum stochastic gradient descent (SGD) methods in the minimization of nonsmooth nonconvex functions, especially in training nonsmooth neural networks. Our framework adaptively generates learning rates based on the historical data of stochastic subgradients and iterates. Under mild conditions, we prove that our proposed framework enjoys global convergence to the stationary points of the objective function in the sense of the conservative field, hence providing convergence guarantees for training nonsmooth neural networks. Based on our proposed framework, we propose a novel learning-rate-free momentum SGD method (LFM). Preliminary numerical experiments reveal that LFM performs comparably to the state-of-the-art learning-rate-free methods (which have not been shown theoretically to be convergence) across well-known neural network training benchmarks.
title Learning-rate-free Momentum SGD with Reshuffling Converges in Nonsmooth Nonconvex Optimization
topic Optimization and Control
url https://arxiv.org/abs/2406.18287