Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Beneventano, Pierfrancesco, Woodworth, Blake |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
por: Phunyaphibarn, Prin, et al.
Publicado: (2023)
por: Phunyaphibarn, Prin, et al.
Publicado: (2023)
On the Trajectories of SGD Without Replacement
por: Beneventano, Pierfrancesco
Publicado: (2023)
por: Beneventano, Pierfrancesco
Publicado: (2023)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
por: Jung, Hyunji, et al.
Publicado: (2025)
por: Jung, Hyunji, et al.
Publicado: (2025)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
por: Andreyev, Arseniy, et al.
Publicado: (2024)
por: Andreyev, Arseniy, et al.
Publicado: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
por: Laus, Hannah, et al.
Publicado: (2025)
por: Laus, Hannah, et al.
Publicado: (2025)
A Local Polyak-Lojasiewicz and Descent Lemma of Gradient Descent For Overparametrized Linear Models
por: Xu, Ziqing, et al.
Publicado: (2025)
por: Xu, Ziqing, et al.
Publicado: (2025)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
por: Cayci, Semih, et al.
Publicado: (2021)
por: Cayci, Semih, et al.
Publicado: (2021)
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
por: Kunstner, Frederik, et al.
Publicado: (2025)
por: Kunstner, Frederik, et al.
Publicado: (2025)
Learning Provably Improves the Convergence of Gradient Descent
por: Song, Qingyu, et al.
Publicado: (2025)
por: Song, Qingyu, et al.
Publicado: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
por: Ward, Rachel, et al.
Publicado: (2023)
por: Ward, Rachel, et al.
Publicado: (2023)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
por: Cayci, Semih, et al.
Publicado: (2024)
por: Cayci, Semih, et al.
Publicado: (2024)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
por: Kassing, Sebastian, et al.
Publicado: (2025)
por: Kassing, Sebastian, et al.
Publicado: (2025)
Finite-Time Analysis of Gradient Descent for Shallow Transformers
por: Arda, Enes, et al.
Publicado: (2026)
por: Arda, Enes, et al.
Publicado: (2026)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
por: Qin, Zhen, et al.
Publicado: (2025)
por: Qin, Zhen, et al.
Publicado: (2025)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
por: Li, Tianyou, et al.
Publicado: (2023)
por: Li, Tianyou, et al.
Publicado: (2023)
Open Problem: Anytime Convergence Rate of Gradient Descent
por: Kornowski, Guy, et al.
Publicado: (2024)
por: Kornowski, Guy, et al.
Publicado: (2024)
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks
por: Gao, Yihang, et al.
Publicado: (2024)
por: Gao, Yihang, et al.
Publicado: (2024)
GANs as Gradient Flows that Converge
por: Huang, Yu-Jui, et al.
Publicado: (2022)
por: Huang, Yu-Jui, et al.
Publicado: (2022)
Does Weight Decay Enhance Training Stability?
por: Saether, Marius, et al.
Publicado: (2026)
por: Saether, Marius, et al.
Publicado: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
por: Xu, Yizhou, et al.
Publicado: (2026)
por: Xu, Yizhou, et al.
Publicado: (2026)
Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence
por: Datar, Adwait, et al.
Publicado: (2025)
por: Datar, Adwait, et al.
Publicado: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
por: Kale, Sacchit, et al.
Publicado: (2026)
por: Kale, Sacchit, et al.
Publicado: (2026)
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation
por: Mishkin, Aaron, et al.
Publicado: (2024)
por: Mishkin, Aaron, et al.
Publicado: (2024)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
por: Ma, Jianhao, et al.
Publicado: (2024)
por: Ma, Jianhao, et al.
Publicado: (2024)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
por: Kong, Boao, et al.
Publicado: (2026)
por: Kong, Boao, et al.
Publicado: (2026)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
por: Xu, Xianliang, et al.
Publicado: (2024)
por: Xu, Xianliang, et al.
Publicado: (2024)
Policy Gradient Converges to the Globally Optimal Policy for Nearly Linear-Quadratic Regulators
por: Han, Yinbin, et al.
Publicado: (2023)
por: Han, Yinbin, et al.
Publicado: (2023)
Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
por: Carmona, René, et al.
Publicado: (2019)
por: Carmona, René, et al.
Publicado: (2019)
Robustness of Iteratively Pre-Conditioned Gradient-Descent Method: The Case of Distributed Linear Regression Problem
por: Chakrabarti, Kushal, et al.
Publicado: (2021)
por: Chakrabarti, Kushal, et al.
Publicado: (2021)
Iterative Pre-Conditioning for Expediting the Gradient-Descent Method: The Distributed Linear Least-Squares Problem
por: Chakrabarti, Kushal, et al.
Publicado: (2020)
por: Chakrabarti, Kushal, et al.
Publicado: (2020)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
por: Oowada, Kanata, et al.
Publicado: (2025)
por: Oowada, Kanata, et al.
Publicado: (2025)
Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares
por: MacDonald, Lachlan Ewen, et al.
Publicado: (2025)
por: MacDonald, Lachlan Ewen, et al.
Publicado: (2025)
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees
por: Maros, Marie, et al.
Publicado: (2022)
por: Maros, Marie, et al.
Publicado: (2022)
Corner Gradient Descent
por: Yarotsky, Dmitry
Publicado: (2025)
por: Yarotsky, Dmitry
Publicado: (2025)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
por: Min, Hancheng, et al.
Publicado: (2025)
por: Min, Hancheng, et al.
Publicado: (2025)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
por: Das, Mohua, et al.
Publicado: (2026)
por: Das, Mohua, et al.
Publicado: (2026)
Stochastic Adaptive Gradient Descent Without Descent
por: Aujol, Jean-François, et al.
Publicado: (2025)
por: Aujol, Jean-François, et al.
Publicado: (2025)
Adaptive Conditional Gradient Descent
por: Khademi, Abbas, et al.
Publicado: (2025)
por: Khademi, Abbas, et al.
Publicado: (2025)
Ejemplares similares
-
Gradient Descent with Polyak's Momentum Finds Flatter Minima via Large Catapults
por: Phunyaphibarn, Prin, et al.
Publicado: (2023) -
On the Trajectories of SGD Without Replacement
por: Beneventano, Pierfrancesco
Publicado: (2023) -
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
por: Jung, Hyunji, et al.
Publicado: (2025) -
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
por: Andreyev, Arseniy, et al.
Publicado: (2024) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)