Learning Provably Improves the Convergence of Gradient Descent
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Qingyu, Lin, Wei, Xu, Hong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
por: Qin, Zhen, et al.
Publicado: (2025)
por: Qin, Zhen, et al.
Publicado: (2025)
Towards Robust Learning to Optimize with Theoretical Guarantees
por: Song, Qingyu, et al.
Publicado: (2025)
por: Song, Qingyu, et al.
Publicado: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
por: Ward, Rachel, et al.
Publicado: (2023)
por: Ward, Rachel, et al.
Publicado: (2023)
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
por: Vaswani, Sharan, et al.
Publicado: (2025)
por: Vaswani, Sharan, et al.
Publicado: (2025)
Unraveling the Gradient Descent Dynamics of Transformers
por: Song, Bingqing, et al.
Publicado: (2024)
por: Song, Bingqing, et al.
Publicado: (2024)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
por: Li, Tianyou, et al.
Publicado: (2023)
por: Li, Tianyou, et al.
Publicado: (2023)
Open Problem: Anytime Convergence Rate of Gradient Descent
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
por: MacDonald, Lachlan Ewen, et al.
Publicado: (2025)
por: MacDonald, Lachlan Ewen, et al.
Publicado: (2025)
Provably Faster Gradient Descent via Long Steps
por: Grimmer, Benjamin
Publicado: (2023)
por: Grimmer, Benjamin
Publicado: (2023)
Mitigating Gradient Bias in Multi-objective Learning: A Provably Convergent Stochastic Approach
por: Fernando, Heshan, et al.
Publicado: (2022)
por: Fernando, Heshan, et al.
Publicado: (2022)
Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence
por: Datar, Adwait, et al.
Publicado: (2025)
por: Datar, Adwait, et al.
Publicado: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
por: Jung, Hyunji, et al.
Publicado: (2025)
por: Jung, Hyunji, et al.
Publicado: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
por: Kale, Sacchit, et al.
Publicado: (2026)
por: Kale, Sacchit, et al.
Publicado: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
por: Mishkin, Aaron, et al.
Publicado: (2024)
por: Mishkin, Aaron, et al.
Publicado: (2024)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
por: Ma, Jianhao, et al.
Publicado: (2024)
por: Ma, Jianhao, et al.
Publicado: (2024)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
por: Kong, Boao, et al.
Publicado: (2026)
por: Kong, Boao, et al.
Publicado: (2026)
Provably Convergent Federated Trilevel Learning
por: Jiao, Yang, et al.
Publicado: (2023)
por: Jiao, Yang, et al.
Publicado: (2023)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
por: Xu, Xianliang, et al.
Publicado: (2024)
por: Xu, Xianliang, et al.
Publicado: (2024)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
por: Chu, Ya-Chi, et al.
Publicado: (2025)
por: Chu, Ya-Chi, et al.
Publicado: (2025)
A New Convergence Analysis of Plug-and-Play Proximal Gradient Descent Under Prior Mismatch
por: Xu, Guixian, et al.
Publicado: (2026)
por: Xu, Guixian, et al.
Publicado: (2026)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
por: Oowada, Kanata, et al.
Publicado: (2025)
por: Oowada, Kanata, et al.
Publicado: (2025)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
por: Cayci, Semih, et al.
Publicado: (2024)
por: Cayci, Semih, et al.
Publicado: (2024)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
por: Beneventano, Pierfrancesco, et al.
Publicado: (2025)
por: Beneventano, Pierfrancesco, et al.
Publicado: (2025)
Convergence of Spectral Descent for Non-smooth Optimization
por: Yang, Yixuan, et al.
Publicado: (2026)
por: Yang, Yixuan, et al.
Publicado: (2026)
On the Convergence of Policy in Unregularized Policy Mirror Descent
por: Lin, Dachao, et al.
Publicado: (2022)
por: Lin, Dachao, et al.
Publicado: (2022)
A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
por: Chu, Tianshu, et al.
Publicado: (2025)
por: Chu, Tianshu, et al.
Publicado: (2025)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
por: Xiao, Minheng, et al.
Publicado: (2024)
por: Xiao, Minheng, et al.
Publicado: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
por: Gao, Yihang, et al.
Publicado: (2024)
por: Gao, Yihang, et al.
Publicado: (2024)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
por: Kassing, Sebastian, et al.
Publicado: (2025)
por: Kassing, Sebastian, et al.
Publicado: (2025)
On the Convergence of Policy Mirror Descent with Temporal Difference Evaluation
por: Liu, Jiacai, et al.
Publicado: (2025)
por: Liu, Jiacai, et al.
Publicado: (2025)
Stochastic Adaptive Gradient Descent Without Descent
por: Aujol, Jean-François, et al.
Publicado: (2025)
por: Aujol, Jean-François, et al.
Publicado: (2025)
Enhancing Fractional Gradient Descent with Learned Optimizers
por: Sobotka, Jan, et al.
Publicado: (2025)
por: Sobotka, Jan, et al.
Publicado: (2025)
Corner Gradient Descent
por: Yarotsky, Dmitry
Publicado: (2025)
por: Yarotsky, Dmitry
Publicado: (2025)
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
On Gradient Descent Ascent for Nonconvex-Concave Minimax Problems
por: Lin, Tianyi, et al.
Publicado: (2019)
por: Lin, Tianyi, et al.
Publicado: (2019)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
por: Xu, Ziqing, et al.
Publicado: (2025)
por: Xu, Ziqing, et al.
Publicado: (2025)
Gauss-Newton Natural Gradient Descent for Shape Learning
por: King, James, et al.
Publicado: (2026)
por: King, James, et al.
Publicado: (2026)
Quantitative Convergence Analysis of Projected Stochastic Gradient Descent for Non-Convex Losses via the Goldstein Subdifferential
por: Zheng, Yuping, et al.
Publicado: (2025)
por: Zheng, Yuping, et al.
Publicado: (2025)
MGDA Converges under Generalized Smoothness, Provably
por: Zhang, Qi, et al.
Publicado: (2024)
por: Zhang, Qi, et al.
Publicado: (2024)
Ejemplares similares
-
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
por: Qin, Zhen, et al.
Publicado: (2025) -
Towards Robust Learning to Optimize with Theoretical Guarantees
por: Song, Qingyu, et al.
Publicado: (2025) -
Convergence of Alternating Gradient Descent for Matrix Factorization
por: Ward, Rachel, et al.
Publicado: (2023) -
Armijo Line-search Can Make (Stochastic) Gradient Descent Provably Faster
por: Vaswani, Sharan, et al.
Publicado: (2025) -
Unraveling the Gradient Descent Dynamics of Transformers
por: Song, Bingqing, et al.
Publicado: (2024)