Learning Provably Improves the Convergence of Gradient Descent

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Song, Qingyu, Lin, Wei, Xu, Hong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908729178324992
author Song, Qingyu
Lin, Wei
Xu, Hong
author_facet Song, Qingyu
Lin, Wei
Xu, Hong
contents Learn to Optimize (L2O) trains deep neural network-based solvers for optimization, achieving success in accelerating convex problems and improving non-convex solutions. However, L2O lacks rigorous theoretical backing for its own training convergence, as existing analyses often use unrealistic assumptions -- a gap this work highlights empirically. We bridge this gap by proving the training convergence of L2O models that learn Gradient Descent (GD) hyperparameters for quadratic programming, leveraging the Neural Tangent Kernel (NTK) theory. We propose a deterministic initialization strategy to support our theoretical results and promote stable training over extended optimization horizons by mitigating gradient explosion. Our L2O framework demonstrates over 50% better optimality than GD and superior robustness over state-of-the-art L2O methods on synthetic datasets. The code of our method can be found from https://github.com/NetX-lab/MathL2OProof-Official.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18092
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Provably Improves the Convergence of Gradient Descent
Song, Qingyu
Lin, Wei
Xu, Hong
Machine Learning
Optimization and Control
Learn to Optimize (L2O) trains deep neural network-based solvers for optimization, achieving success in accelerating convex problems and improving non-convex solutions. However, L2O lacks rigorous theoretical backing for its own training convergence, as existing analyses often use unrealistic assumptions -- a gap this work highlights empirically. We bridge this gap by proving the training convergence of L2O models that learn Gradient Descent (GD) hyperparameters for quadratic programming, leveraging the Neural Tangent Kernel (NTK) theory. We propose a deterministic initialization strategy to support our theoretical results and promote stable training over extended optimization horizons by mitigating gradient explosion. Our L2O framework demonstrates over 50% better optimality than GD and superior robustness over state-of-the-art L2O methods on synthetic datasets. The code of our method can be found from https://github.com/NetX-lab/MathL2OProof-Official.
title Learning Provably Improves the Convergence of Gradient Descent
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2501.18092