Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beneventano, Pierfrancesco, Woodworth, Blake |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2023)
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2023)
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
von: Jung, Hyunji, et al.
Veröffentlicht: (2025)
von: Jung, Hyunji, et al.
Veröffentlicht: (2025)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
von: Xu, Ziqing, et al.
Veröffentlicht: (2025)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
von: Cayci, Semih, et al.
Veröffentlicht: (2021)
von: Cayci, Semih, et al.
Veröffentlicht: (2021)
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
von: Kunstner, Frederik, et al.
Veröffentlicht: (2025)
von: Kunstner, Frederik, et al.
Veröffentlicht: (2025)
Learning Provably Improves the Convergence of Gradient Descent
von: Song, Qingyu, et al.
Veröffentlicht: (2025)
von: Song, Qingyu, et al.
Veröffentlicht: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
von: Ward, Rachel, et al.
Veröffentlicht: (2023)
von: Ward, Rachel, et al.
Veröffentlicht: (2023)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
von: Kassing, Sebastian, et al.
Veröffentlicht: (2025)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
von: Arda, Enes, et al.
Veröffentlicht: (2026)
von: Arda, Enes, et al.
Veröffentlicht: (2026)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
von: Qin, Zhen, et al.
Veröffentlicht: (2025)
von: Qin, Zhen, et al.
Veröffentlicht: (2025)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
von: Li, Tianyou, et al.
Veröffentlicht: (2023)
von: Li, Tianyou, et al.
Veröffentlicht: (2023)
Open Problem: Anytime Convergence Rate of Gradient Descent
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
von: Kornowski, Guy, et al.
Veröffentlicht: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
von: Gao, Yihang, et al.
Veröffentlicht: (2024)
GANs as Gradient Flows that Converge
von: Huang, Yu-Jui, et al.
Veröffentlicht: (2022)
von: Huang, Yu-Jui, et al.
Veröffentlicht: (2022)
Does Weight Decay Enhance Training Stability?
von: Saether, Marius, et al.
Veröffentlicht: (2026)
von: Saether, Marius, et al.
Veröffentlicht: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
von: Xu, Yizhou, et al.
Veröffentlicht: (2026)
Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence
von: Datar, Adwait, et al.
Veröffentlicht: (2025)
von: Datar, Adwait, et al.
Veröffentlicht: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
von: Ma, Jianhao, et al.
Veröffentlicht: (2024)
von: Ma, Jianhao, et al.
Veröffentlicht: (2024)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
von: Kong, Boao, et al.
Veröffentlicht: (2026)
von: Kong, Boao, et al.
Veröffentlicht: (2026)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
von: Han, Yinbin, et al.
Veröffentlicht: (2023)
von: Han, Yinbin, et al.
Veröffentlicht: (2023)
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
von: Carmona, René, et al.
Veröffentlicht: (2019)
von: Carmona, René, et al.
Veröffentlicht: (2019)
Robustness of Iteratively Pre-Conditioned Gradient-Descent Method: The Case of Distributed Linear Regression Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
Iterative Pre-Conditioning for Expediting the Gradient-Descent Method: The Distributed Linear Least-Squares Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
von: Oowada, Kanata, et al.
Veröffentlicht: (2025)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
von: MacDonald, Lachlan Ewen, et al.
Veröffentlicht: (2025)
von: MacDonald, Lachlan Ewen, et al.
Veröffentlicht: (2025)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees
von: Maros, Marie, et al.
Veröffentlicht: (2022)
von: Maros, Marie, et al.
Veröffentlicht: (2022)
Corner Gradient Descent
von: Yarotsky, Dmitry
Veröffentlicht: (2025)
von: Yarotsky, Dmitry
Veröffentlicht: (2025)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
von: Das, Mohua, et al.
Veröffentlicht: (2026)
von: Das, Mohua, et al.
Veröffentlicht: (2026)
Stochastic Adaptive Gradient Descent Without Descent
von: Aujol, Jean-François, et al.
Veröffentlicht: (2025)
von: Aujol, Jean-François, et al.
Veröffentlicht: (2025)
Adaptive Conditional Gradient Descent
von: Khademi, Abbas, et al.
Veröffentlicht: (2025)
von: Khademi, Abbas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2023) -
On the Trajectories of SGD Without Replacement
von: Beneventano, Pierfrancesco
Veröffentlicht: (2023) -
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
von: Jung, Hyunji, et al.
Veröffentlicht: (2025) -
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
von: Andreyev, Arseniy, et al.
Veröffentlicht: (2024) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
von: Beneventano, Pierfrancesco, et al.
Veröffentlicht: (2024)