Convergence, Sticking and Escape: Stochastic Dynamics Near Critical Points in SGD
Fuente:
arXiv
Saved in:
| Main Authors: | Dudukalov, Dmitry, Logachov, Artem, Lotov, Vladimir, Prasolov, Timofei, Prokopenko, Evgeny, Tarasenko, Anton |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Basic model for ranking microfinance institutions
by: Dudukalov, Dmitry, et al.
Published: (2025)
by: Dudukalov, Dmitry, et al.
Published: (2025)
Central Limit Theorem on Symmetric Kullback-Leibler (KL) Divergence
by: Rojas, Helder, et al.
Published: (2024)
by: Rojas, Helder, et al.
Published: (2024)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
by: Chen, Zaiwei, et al.
Published: (2026)
by: Chen, Zaiwei, et al.
Published: (2026)
Error estimates between SGD with momentum and underdamped Langevin diffusion
by: Guillin, Arnaud, et al.
Published: (2024)
by: Guillin, Arnaud, et al.
Published: (2024)
Effective continuous equations for adaptive SGD: a stochastic analysis view
by: Callisti, Luca, et al.
Published: (2025)
by: Callisti, Luca, et al.
Published: (2025)
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD
by: Wan, Yijun, et al.
Published: (2023)
by: Wan, Yijun, et al.
Published: (2023)
Convergent Stochastic Training of Attention and Understanding LoRA
by: Sun, Zhengkai, et al.
Published: (2026)
by: Sun, Zhengkai, et al.
Published: (2026)
Decentralized Proximal Stochastic Gradient Langevin Dynamics
by: Islam, Mohammad Rafiqul, et al.
Published: (2026)
by: Islam, Mohammad Rafiqul, et al.
Published: (2026)
Weak Convergence Analysis of Online Neural Actor-Critic Algorithms
by: Lam, Samuel Chun-Hei, et al.
Published: (2024)
by: Lam, Samuel Chun-Hei, et al.
Published: (2024)
Minima and Critical Points of the Bethe Free Energy Are Invariant Under Deformation Retractions of Factor Graphs
by: Sergeant-Perthuis, Grégoire, et al.
Published: (2025)
by: Sergeant-Perthuis, Grégoire, et al.
Published: (2025)
Controlling the Flow: Stability and Convergence for Stochastic Gradient Descent with Decaying Regularization
by: Kassing, Sebastian, et al.
Published: (2025)
by: Kassing, Sebastian, et al.
Published: (2025)
ResNets of All Shapes and Sizes: Convergence of Training Dynamics in the Large-scale Limit
by: Chaintron, Louis-Pierre, et al.
Published: (2026)
by: Chaintron, Louis-Pierre, et al.
Published: (2026)
Convergence of Two Time-Scale Stochastic Approximation: A Martingale Approach
by: Vidyasagar, Mathukumalli
Published: (2026)
by: Vidyasagar, Mathukumalli
Published: (2026)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
by: Fouque, Jean-Pierre, et al.
Published: (2025)
by: Fouque, Jean-Pierre, et al.
Published: (2025)
From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes
by: Chen, Zaiwei, et al.
Published: (2025)
by: Chen, Zaiwei, et al.
Published: (2025)
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent
by: Schertzer, Adrien, et al.
Published: (2024)
by: Schertzer, Adrien, et al.
Published: (2024)
Flatness-Aware Stochastic Gradient Langevin Dynamics
by: Bruno, Stefano, et al.
Published: (2025)
by: Bruno, Stefano, et al.
Published: (2025)
Exponential Convergence Guarantees for Iterative Markovian Fitting
by: Silveri, Marta Gentiloni, et al.
Published: (2025)
by: Silveri, Marta Gentiloni, et al.
Published: (2025)
Stochastic Operator Network: A Stochastic Maximum Principle Based Approach to Operator Learning
by: Bausback, Ryan, et al.
Published: (2025)
by: Bausback, Ryan, et al.
Published: (2025)
Riemannian Langevin Dynamics: Strong Convergence of Geometric Euler-Maruyama Scheme
by: Zhan, Zhiyuan, et al.
Published: (2026)
by: Zhan, Zhiyuan, et al.
Published: (2026)
Deep Learning for Computing Convergence Rates of Markov Chains
by: Qu, Yanlin, et al.
Published: (2024)
by: Qu, Yanlin, et al.
Published: (2024)
Convergence Analysis of Newton's Method for Neural Networks in the Overparameterized Limit
by: Riedl, Konstantin, et al.
Published: (2026)
by: Riedl, Konstantin, et al.
Published: (2026)
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
by: Armacki, Aleksandar, et al.
Published: (2024)
by: Armacki, Aleksandar, et al.
Published: (2024)
Characterizing Dynamical Stability of Stochastic Gradient Descent in Overparameterized Learning
by: Chemnitz, Dennis, et al.
Published: (2024)
by: Chemnitz, Dennis, et al.
Published: (2024)
Learning the Infinitesimal Generator of Stochastic Diffusion Processes
by: Kostic, Vladimir R., et al.
Published: (2024)
by: Kostic, Vladimir R., et al.
Published: (2024)
Structural and Convergence Analysis of Discrete-Time Denoising Diffusion Probabilistic Models
by: Nakano, Yumiharu
Published: (2024)
by: Nakano, Yumiharu
Published: (2024)
Random-Bridges as Stochastic Transports for Generative Models
by: Goria, Stefano, et al.
Published: (2025)
by: Goria, Stefano, et al.
Published: (2025)
Convergence Analysis for General Probability Flow ODEs of Diffusion Models in Wasserstein Distances
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Wasserstein Convergence Guarantees for a General Class of Score-Based Generative Models
by: Gao, Xuefeng, et al.
Published: (2023)
by: Gao, Xuefeng, et al.
Published: (2023)
Limit Theorems for Stochastic Gradient Descent with Infinite Variance
by: Blanchet, Jose, et al.
Published: (2024)
by: Blanchet, Jose, et al.
Published: (2024)
Partially Stochastic Infinitely Deep Bayesian Neural Networks
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
Approximating G(t)/GI/1 queues with deep learning
by: Sherzer, Eliran, et al.
Published: (2024)
by: Sherzer, Eliran, et al.
Published: (2024)
Stabilizing Temporal Difference Learning via Implicit Stochastic Recursion
by: Kim, Hwanwoo, et al.
Published: (2025)
by: Kim, Hwanwoo, et al.
Published: (2025)
Flow Matching: Markov Kernels, Stochastic Processes and Transport Plans
by: Wald, Christian, et al.
Published: (2025)
by: Wald, Christian, et al.
Published: (2025)
Approximation to Deep Q-Network by Stochastic Delay Differential Equations
by: Lu, Jianya, et al.
Published: (2025)
by: Lu, Jianya, et al.
Published: (2025)
Stochastic Scaling Limits and Synchronization by Noise in Deep Transformer Models
by: Agazzi, Andrea, et al.
Published: (2026)
by: Agazzi, Andrea, et al.
Published: (2026)
Convergence Error Analysis of Reflected Gradient Langevin Dynamics for Globally Optimizing Non-Convex Constrained Problems
by: Sato, Kanji, et al.
Published: (2022)
by: Sato, Kanji, et al.
Published: (2022)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
by: Agrawalla, Bhavya, et al.
Published: (2023)
by: Agrawalla, Bhavya, et al.
Published: (2023)
Convergence of Unadjusted Langevin in High Dimensions: Delocalization of Bias
by: Chen, Yifan, et al.
Published: (2024)
by: Chen, Yifan, et al.
Published: (2024)
Similar Items
-
Basic model for ranking microfinance institutions
by: Dudukalov, Dmitry, et al.
Published: (2025) -
Central Limit Theorem on Symmetric Kullback-Leibler (KL) Divergence
by: Rojas, Helder, et al.
Published: (2024) -
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023) -
Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework
by: Chen, Zaiwei, et al.
Published: (2026) -
Error estimates between SGD with momentum and underdamped Langevin diffusion
by: Guillin, Arnaud, et al.
Published: (2024)