Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Shuang, Montúfar, Guido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression
by: Liang, Shuang, et al.
Published: (2024)
by: Liang, Shuang, et al.
Published: (2024)
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
by: Tyurin, Alexander
Published: (2024)
by: Tyurin, Alexander
Published: (2024)
On the Local Complexity of Linear Regions in Deep ReLU Networks
by: Patel, Niket, et al.
Published: (2024)
by: Patel, Niket, et al.
Published: (2024)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
by: Meng, Si Yi, et al.
Published: (2025)
by: Meng, Si Yi, et al.
Published: (2025)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Low Rank Gradients and Where to Find Them
by: Sonthalia, Rishi, et al.
Published: (2025)
by: Sonthalia, Rishi, et al.
Published: (2025)
On the Convergence of Gradient Descent for Large Learning Rates
by: Crăciun, Alexandru, et al.
Published: (2024)
by: Crăciun, Alexandru, et al.
Published: (2024)
On the Convergence Rate of LoRA Gradient Descent
by: Mu, Siqiao, et al.
Published: (2025)
by: Mu, Siqiao, et al.
Published: (2025)
Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descent
by: Hennick, Max, et al.
Published: (2025)
by: Hennick, Max, et al.
Published: (2025)
Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning
by: Jacobs, Tom, et al.
Published: (2026)
by: Jacobs, Tom, et al.
Published: (2026)
Accelerated Gradient Descent for Faster Convergence with Minimal Overhead
by: Graca, Manuel, et al.
Published: (2026)
by: Graca, Manuel, et al.
Published: (2026)
Stochastic Gradient Descent in Non-Convex Problems: Asymptotic Convergence with Relaxed Step-Size via Stopping Time Methods
by: Jin, Ruinan, et al.
Published: (2025)
by: Jin, Ruinan, et al.
Published: (2025)
Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
by: Sung, Kyle, et al.
Published: (2025)
by: Sung, Kyle, et al.
Published: (2025)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
by: Zhang, Ruiqi, et al.
Published: (2025)
by: Zhang, Ruiqi, et al.
Published: (2025)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
by: Ward, Rachel, et al.
Published: (2023)
by: Ward, Rachel, et al.
Published: (2023)
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees
by: Wu, Yongtao, et al.
Published: (2025)
by: Wu, Yongtao, et al.
Published: (2025)
Constraining the outputs of ReLU neural networks
by: Alexandr, Yulia, et al.
Published: (2025)
by: Alexandr, Yulia, et al.
Published: (2025)
Understanding Learning Invariance in Deep Linear Networks
by: Duan, Hao, et al.
Published: (2025)
by: Duan, Hao, et al.
Published: (2025)
Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability
by: Gan, Eric
Published: (2026)
by: Gan, Eric
Published: (2026)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
by: Qin, Zhen, et al.
Published: (2025)
by: Qin, Zhen, et al.
Published: (2025)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
by: Li, Tianyou, et al.
Published: (2023)
by: Li, Tianyou, et al.
Published: (2023)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Accelerating Convergence of Stein Variational Gradient Descent via Deep Unfolding
by: Kawamura, Yuya, et al.
Published: (2024)
by: Kawamura, Yuya, et al.
Published: (2024)
Convergence Analysis of Fractional Gradient Descent
by: Aggarwal, Ashwani
Published: (2023)
by: Aggarwal, Ashwani
Published: (2023)
Bounds for the smallest eigenvalue of the NTK for arbitrary spherical data of arbitrary dimension
by: Karhadkar, Kedar, et al.
Published: (2024)
by: Karhadkar, Kedar, et al.
Published: (2024)
Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence
by: Datar, Adwait, et al.
Published: (2025)
by: Datar, Adwait, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
by: Ma, Jianhao, et al.
Published: (2024)
by: Ma, Jianhao, et al.
Published: (2024)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Enhancing Policy Gradient with the Polyak Step-Size Adaption
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Provably Faster Gradient Descent via Long Steps
by: Grimmer, Benjamin
Published: (2023)
by: Grimmer, Benjamin
Published: (2023)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Similar Items
-
Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression
by: Liang, Shuang, et al.
Published: (2024) -
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
by: Tyurin, Alexander
Published: (2024) -
On the Local Complexity of Linear Regions in Deep ReLU Networks
by: Patel, Niket, et al.
Published: (2024) -
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023) -
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)