On the Convergence of Gradient Descent for Large Learning Rates
Fuente:
arXiv
Saved in:
| Main Authors: | Crăciun, Alexandru, Ghoshdastidar, Debarghya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025)
by: Crăciun, Alexandru, et al.
Published: (2025)
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
by: Engquist, Björn, et al.
Published: (2022)
by: Engquist, Björn, et al.
Published: (2022)
Stochastic Gradient Descent Revisited
by: Louzi, Azar
Published: (2024)
by: Louzi, Azar
Published: (2024)
A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties
by: Qiu, Junwen, et al.
Published: (2023)
by: Qiu, Junwen, et al.
Published: (2023)
Bilevel Learning via Inexact Stochastic Gradient Descent
by: Salehi, Mohammad Sadegh, et al.
Published: (2025)
by: Salehi, Mohammad Sadegh, et al.
Published: (2025)
Optimal Asymptotic Rates for (Stochastic) Gradient Descent under the Local PL-Condition: A Geometric Approach
by: Kassing, Sebastian, et al.
Published: (2026)
by: Kassing, Sebastian, et al.
Published: (2026)
Clust-Splitter - an Efficient Nonsmooth Optimization-Based Algorithm for Clustering Large Datasets
by: Lampainen, Jenni, et al.
Published: (2025)
by: Lampainen, Jenni, et al.
Published: (2025)
A KL-based Analysis Framework with Applications to Non-Descent Optimization Methods
by: Qiu, Junwen, et al.
Published: (2024)
by: Qiu, Junwen, et al.
Published: (2024)
Shuffling the Stochastic Mirror Descent via Dual Lipschitz Continuity and Kernel Conditioning
by: Qiu, Junwen, et al.
Published: (2026)
by: Qiu, Junwen, et al.
Published: (2026)
Convergence of Two Time-Scale Stochastic Approximation: A Martingale Approach
by: Vidyasagar, Mathukumalli
Published: (2026)
by: Vidyasagar, Mathukumalli
Published: (2026)
Effective Front-Descent Algorithms with Convergence Guarantees
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Policy Gradient Algorithms for Robust MDPs with Non-Rectangular Uncertainty Sets
by: Li, Mengmeng, et al.
Published: (2023)
by: Li, Mengmeng, et al.
Published: (2023)
Communication Compression for Byzantine Robust Learning: New Efficient Algorithms and Improved Rates
by: Rammal, Ahmad, et al.
Published: (2023)
by: Rammal, Ahmad, et al.
Published: (2023)
FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection
by: He, Yutong, et al.
Published: (2026)
by: He, Yutong, et al.
Published: (2026)
Gradient Norm Regularization Second-Order Algorithms for Solving Nonconvex-Strongly Concave Minimax Problems
by: Wang, Jun-Lin, et al.
Published: (2024)
by: Wang, Jun-Lin, et al.
Published: (2024)
Taming Nonconvex Stochastic Mirror Descent with General Bregman Divergence
by: Fatkhullin, Ilyas, et al.
Published: (2024)
by: Fatkhullin, Ilyas, et al.
Published: (2024)
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing
by: Yi, Zeji, et al.
Published: (2026)
by: Yi, Zeji, et al.
Published: (2026)
A Globally Convergent Gradient Method with Momentum
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
A New Kernel Regularity Condition for Distributed Mirror Descent: Broader Coverage and Simpler Analysis
by: Qiu, Junwen, et al.
Published: (2026)
by: Qiu, Junwen, et al.
Published: (2026)
Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints
by: Malo, Pekka, et al.
Published: (2024)
by: Malo, Pekka, et al.
Published: (2024)
Natural Gradient VI: Guarantees for Non-Conjugate Models
by: Sun, Fangyuan, et al.
Published: (2025)
by: Sun, Fangyuan, et al.
Published: (2025)
A Large Deviations Perspective on Policy Gradient Algorithms
by: Jongeneel, Wouter, et al.
Published: (2023)
by: Jongeneel, Wouter, et al.
Published: (2023)
Polyak's Heavy Ball Method Achieves Accelerated Local Rate of Convergence under Polyak-Lojasiewicz Inequality
by: Kassing, Sebastian, et al.
Published: (2024)
by: Kassing, Sebastian, et al.
Published: (2024)
Pareto-optimal Trade-offs Between Communication and Computation with Flexible Gradient Tracking
by: Huang, Yan, et al.
Published: (2025)
by: Huang, Yan, et al.
Published: (2025)
Optimization without Retraction on the Random Generalized Stiefel Manifold
by: Vary, Simon, et al.
Published: (2024)
by: Vary, Simon, et al.
Published: (2024)
A New Random Reshuffling Method for Nonsmooth Nonconvex Finite-sum Optimization
by: Qiu, Junwen, et al.
Published: (2023)
by: Qiu, Junwen, et al.
Published: (2023)
Convergence Conditions for Stochastic Line Search Based Optimization of Over-parametrized Models
by: Lapucci, Matteo, et al.
Published: (2024)
by: Lapucci, Matteo, et al.
Published: (2024)
Two trust region type algorithms for solving nonconvex-strongly concave minimax problems
by: Yao, Tongliang, et al.
Published: (2024)
by: Yao, Tongliang, et al.
Published: (2024)
A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems
by: Wang, Junlin, et al.
Published: (2024)
by: Wang, Junlin, et al.
Published: (2024)
Global Solutions to Non-Convex Functional Constrained Problems with Hidden Convexity
by: Fatkhullin, Ilyas, et al.
Published: (2025)
by: Fatkhullin, Ilyas, et al.
Published: (2025)
Exact Convex Reformulations of Linear Neural Networks via Completely Positive Lifting
by: Prakhya, Karthik, et al.
Published: (2026)
by: Prakhya, Karthik, et al.
Published: (2026)
Theoretical Framework for Tempered Fractional Gradient Descent: Application to Breast Cancer Classification
by: Naifar, Omar
Published: (2025)
by: Naifar, Omar
Published: (2025)
SPAM: Stochastic Proximal Point Method with Momentum Variance Reduction for Non-convex Cross-Device Federated Learning
by: Karagulyan, Avetik, et al.
Published: (2024)
by: Karagulyan, Avetik, et al.
Published: (2024)
Det-CGD: Compressed Gradient Descent with Matrix Stepsizes for Non-Convex Optimization
by: Li, Hanmin, et al.
Published: (2023)
by: Li, Hanmin, et al.
Published: (2023)
Alternating minimization for square root principal component pursuit
by: Deng, Shengxiang, et al.
Published: (2024)
by: Deng, Shengxiang, et al.
Published: (2024)
Stochastic First-Order Methods with Non-smooth and Non-Euclidean Proximal Terms for Nonconvex High-Dimensional Stochastic Optimization
by: Xie, Yue, et al.
Published: (2024)
by: Xie, Yue, et al.
Published: (2024)
High Probability Guarantees for Random Reshuffling
by: Yu, Hengxu, et al.
Published: (2023)
by: Yu, Hengxu, et al.
Published: (2023)
Inexact Riemannian Gradient Descent Method for Nonconvex Optimization
by: Zhou, Juan, et al.
Published: (2024)
by: Zhou, Juan, et al.
Published: (2024)
A Generalized Version of Chung's Lemma and its Applications
by: Jiang, Li, et al.
Published: (2024)
by: Jiang, Li, et al.
Published: (2024)
IGT-OMD: Implicit Gradient Transport for Decision-Focused Learning under Delayed Feedback
by: Amoh, Benjamin, et al.
Published: (2026)
by: Amoh, Benjamin, et al.
Published: (2026)
Similar Items
-
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025) -
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
by: Engquist, Björn, et al.
Published: (2022) -
Stochastic Gradient Descent Revisited
by: Louzi, Azar
Published: (2024) -
A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties
by: Qiu, Junwen, et al.
Published: (2023) -
Bilevel Learning via Inexact Stochastic Gradient Descent
by: Salehi, Mohammad Sadegh, et al.
Published: (2025)