On the global convergence of gradient descent for wide shallow models with bounded nonlinearities
Fuente:
arXiv
Saved in:
| Main Authors: | Petit, Romain, Poon, Clarice, Peyré, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024)
by: Barboni, Raphaël, et al.
Published: (2024)
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Curvature of optimal transport with respect to the cost and applications to inverse optimal transport
by: Peyré, Gabriel, et al.
Published: (2026)
by: Peyré, Gabriel, et al.
Published: (2026)
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024)
by: Jentzen, Arnulf, et al.
Published: (2024)
Global convergence of gradient descent for phase retrieval
by: Fougereux, Théodore, et al.
Published: (2024)
by: Fougereux, Théodore, et al.
Published: (2024)
Muon Dynamics as a Spectral Wasserstein Flow
by: Peyré, Gabriel
Published: (2026)
by: Peyré, Gabriel
Published: (2026)
Optimal and Diffusion Transports in Machine Learning
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Optimal Transport for Machine Learners
by: Peyré, Gabriel
Published: (2025)
by: Peyré, Gabriel
Published: (2025)
Problem-dependent convergence bounds for randomized linear gradient compression
by: Flynn, Thomas, et al.
Published: (2024)
by: Flynn, Thomas, et al.
Published: (2024)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
by: Davis, Damek, et al.
Published: (2026)
by: Davis, Damek, et al.
Published: (2026)
Abide by the Law and Follow the Flow: Conservation Laws for Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2023)
by: Marcotte, Sibylle, et al.
Published: (2023)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
New logarithmic step size for stochastic gradient descent
by: Shamaee, M. Soheil, et al.
Published: (2024)
by: Shamaee, M. Soheil, et al.
Published: (2024)
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
by: Barboni, Raphaël, et al.
Published: (2025)
by: Barboni, Raphaël, et al.
Published: (2025)
Linear convergence of proximal descent schemes on the Wasserstein space
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
by: Lascu, Razvan-Andrei, et al.
Published: (2024)
Error bounds for particle gradient descent, and extensions of the log-Sobolev and Talagrand inequalities
by: Caprio, Rocco, et al.
Published: (2024)
by: Caprio, Rocco, et al.
Published: (2024)
The duality structure gradient descent algorithm: analysis and applications to neural networks
by: Flynn, Thomas
Published: (2017)
by: Flynn, Thomas
Published: (2017)
On the stability of gradient descent with second order dynamics for time-varying cost functions
by: Gibson, Travis E., et al.
Published: (2024)
by: Gibson, Travis E., et al.
Published: (2024)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
by: Lugosi, Gabor, et al.
Published: (2024)
by: Lugosi, Gabor, et al.
Published: (2024)
A stochastic gradient descent algorithm with random search directions
by: Gbaguidi, Eméric
Published: (2025)
by: Gbaguidi, Eméric
Published: (2025)
Error dynamics of mini-batch gradient descent with random reshuffling for least squares regression
by: Lok, Jackie, et al.
Published: (2024)
by: Lok, Jackie, et al.
Published: (2024)
Almost sure convergence rates of stochastic gradient methods under gradient domination
by: Weissmann, Simon, et al.
Published: (2024)
by: Weissmann, Simon, et al.
Published: (2024)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
by: An, Jing, et al.
Published: (2023)
by: An, Jing, et al.
Published: (2023)
Convergence of flow-based generative models via proximal gradient descent in Wasserstein space
by: Cheng, Xiuyuan, et al.
Published: (2023)
by: Cheng, Xiuyuan, et al.
Published: (2023)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Local linear convergence of gradient methods for overparameterized Gaussian mixtures
by: Wang, Jingxing, et al.
Published: (2026)
by: Wang, Jingxing, et al.
Published: (2026)
On the non-convexity issue in the radial Calderón problem
by: Alberti, Giovanni S., et al.
Published: (2025)
by: Alberti, Giovanni S., et al.
Published: (2025)
Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process
by: De Castro, Yohann, et al.
Published: (2026)
by: De Castro, Yohann, et al.
Published: (2026)
When majority rules, minority loses: bias amplification of gradient descent
by: Bachoc, François, et al.
Published: (2025)
by: Bachoc, François, et al.
Published: (2025)
Manifold constrained steepest descent
by: Yang, Kaiwei, et al.
Published: (2026)
by: Yang, Kaiwei, et al.
Published: (2026)
On the convergence analysis of the decentralized projected gradient descent method
by: Choi, Woocheol, et al.
Published: (2023)
by: Choi, Woocheol, et al.
Published: (2023)
Flattened one-bit stochastic gradient descent: compressed distributed optimization with controlled variance
by: Stollenwerk, Alexander, et al.
Published: (2024)
by: Stollenwerk, Alexander, et al.
Published: (2024)
On diffusion-based generative models and their error bounds: The log-concave case with full convergence estimates
by: Bruno, Stefano, et al.
Published: (2023)
by: Bruno, Stefano, et al.
Published: (2023)
Long-time dynamics and universality of nonconvex gradient descent
by: Han, Qiyang
Published: (2025)
by: Han, Qiyang
Published: (2025)
From Score Matching to Diffusion: A Fine-Grained Error Analysis in the Gaussian Setting
by: Hurault, Samuel, et al.
Published: (2025)
by: Hurault, Samuel, et al.
Published: (2025)
Convergence of two-timescale gradient descent ascent dynamics: finite-dimensional and mean-field perspectives
by: An, Jing, et al.
Published: (2025)
by: An, Jing, et al.
Published: (2025)
Training Infinitely Deep and Wide Transformers
by: Barboni, Raphaël, et al.
Published: (2026)
by: Barboni, Raphaël, et al.
Published: (2026)
Proximal gradient-type method with generalized distance and convergence analysis without global descent lemma
by: Yagishita, Shotaro, et al.
Published: (2025)
by: Yagishita, Shotaro, et al.
Published: (2025)
Learning mirror maps in policy mirror descent
by: Alfano, Carlo, et al.
Published: (2024)
by: Alfano, Carlo, et al.
Published: (2024)
Riemannian coordinate descent algorithms on matrix manifolds
by: Han, Andi, et al.
Published: (2024)
by: Han, Andi, et al.
Published: (2024)
Similar Items
-
Understanding the training of infinitely deep and wide ResNets with Conditional Optimal Transport
by: Barboni, Raphaël, et al.
Published: (2024) -
Robust Sublinear Convergence Rates for Iterative Bregman Projections
by: Peyré, Gabriel
Published: (2026) -
Curvature of optimal transport with respect to the cost and applications to inverse optimal transport
by: Peyré, Gabriel, et al.
Published: (2026) -
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks
by: Jentzen, Arnulf, et al.
Published: (2024) -
Global convergence of gradient descent for phase retrieval
by: Fougereux, Théodore, et al.
Published: (2024)