Anytime Acceleration of Gradient Descent

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Zihan, Lee, Jason D., Du, Simon S., Chen, Yuxin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929619454656512
author Zhang, Zihan
Lee, Jason D.
Du, Simon S.
Chen, Yuxin
author_facet Zhang, Zihan
Lee, Jason D.
Du, Simon S.
Chen, Yuxin
contents This work investigates stepsize-based acceleration of gradient descent with {\em anytime} convergence guarantees. For smooth (non-strongly) convex optimization, we propose a stepsize schedule that allows gradient descent to achieve convergence guarantees of $O(T^{-1.119})$ for any stopping time $T$, where the stepsize schedule is predetermined without prior knowledge of the stopping time. This result provides an affirmative answer to a COLT open problem \citep{kornowski2024open} regarding whether stepsize-based acceleration can yield anytime convergence rates of $o(T^{-1})$. We further extend our theory to yield anytime convergence guarantees of $\exp(-Ω(T/κ^{0.893}))$ for smooth and strongly convex optimization, with $κ$ being the condition number.
format Preprint
id arxiv_https___arxiv_org_abs_2411_17668
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Anytime Acceleration of Gradient Descent
Zhang, Zihan
Lee, Jason D.
Du, Simon S.
Chen, Yuxin
Machine Learning
Systems and Control
Optimization and Control
This work investigates stepsize-based acceleration of gradient descent with {\em anytime} convergence guarantees. For smooth (non-strongly) convex optimization, we propose a stepsize schedule that allows gradient descent to achieve convergence guarantees of $O(T^{-1.119})$ for any stopping time $T$, where the stepsize schedule is predetermined without prior knowledge of the stopping time. This result provides an affirmative answer to a COLT open problem \citep{kornowski2024open} regarding whether stepsize-based acceleration can yield anytime convergence rates of $o(T^{-1})$. We further extend our theory to yield anytime convergence guarantees of $\exp(-Ω(T/κ^{0.893}))$ for smooth and strongly convex optimization, with $κ$ being the condition number.
title Anytime Acceleration of Gradient Descent
topic Machine Learning
Systems and Control
Optimization and Control
url https://arxiv.org/abs/2411.17668