Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ning, Zhang, Yikuan, Ouyang, Qi, Tang, Chao, Tu, Yuhai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
by: Kühn, Marcel, et al.
Published: (2023)
by: Kühn, Marcel, et al.
Published: (2023)
Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems
by: Angelini, Maria Chiara, et al.
Published: (2023)
by: Angelini, Maria Chiara, et al.
Published: (2023)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Random Matrix Theory for Stochastic Gradient Descent
by: Park, Chanju, et al.
Published: (2024)
by: Park, Chanju, et al.
Published: (2024)
Convergence Acceleration of Markov Chain Monte Carlo-based Gradient Descent by Deep Unfolding
by: Hagiwara, Ryo, et al.
Published: (2024)
by: Hagiwara, Ryo, et al.
Published: (2024)
Fundamental limits of learning in sequence multi-index models and deep attention networks: High-dimensional asymptotics and sharp thresholds
by: Troiani, Emanuele, et al.
Published: (2025)
by: Troiani, Emanuele, et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
by: Veiga, Rodrigo, et al.
Published: (2024)
by: Veiga, Rodrigo, et al.
Published: (2024)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Quantum Equilibrium Propagation: Gradient-Descent Training of Quantum Systems
by: Scellier, Benjamin
Published: (2024)
by: Scellier, Benjamin
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Stochastic quantum models for the dynamics of power grids
by: Guichard, Pierrick, et al.
Published: (2024)
by: Guichard, Pierrick, et al.
Published: (2024)
High-dimensional dynamical systems: co-existence of attractors, phase transitions, maximal Lyapunov exponent and response to periodic drive
by: Fournier, Samantha J., et al.
Published: (2025)
by: Fournier, Samantha J., et al.
Published: (2025)
Transient dynamics of associative memory models
by: Clark, David G.
Published: (2025)
by: Clark, David G.
Published: (2025)
Applying statistical learning theory to deep learning
by: Gerbelot, Cédric, et al.
Published: (2023)
by: Gerbelot, Cédric, et al.
Published: (2023)
Factual recall in linear associative memories: sharp asymptotics and mechanistic insights
by: Giorlandino, Alessio, et al.
Published: (2026)
by: Giorlandino, Alessio, et al.
Published: (2026)
Transferable potential for molecular dynamics simulations of borosilicate glasses and structural comparison of machine learning optimized parameters
by: Yang, Kai, et al.
Published: (2025)
by: Yang, Kai, et al.
Published: (2025)
Quantum transport under oscillatory drive with disordered amplitude
by: Tiwari, Vatsana, et al.
Published: (2024)
by: Tiwari, Vatsana, et al.
Published: (2024)
Asymptotic theory of in-context learning by linear attention
by: Lu, Yue M., et al.
Published: (2024)
by: Lu, Yue M., et al.
Published: (2024)
High-dimensional learning of narrow neural networks
by: Cui, Hugo
Published: (2024)
by: Cui, Hugo
Published: (2024)
Stochastic weight matrix dynamics during learning and Dyson Brownian motion
by: Aarts, Gert, et al.
Published: (2024)
by: Aarts, Gert, et al.
Published: (2024)
Representational Drift and Learning-Induced Stabilization in the Olfactory Cortex
by: Morales, Guillermo B., et al.
Published: (2024)
by: Morales, Guillermo B., et al.
Published: (2024)
A unified theory of feature learning in RNNs and DNNs
by: Bauer, Jan P., et al.
Published: (2026)
by: Bauer, Jan P., et al.
Published: (2026)
Dynamics of Meta-learning Representation in the Teacher-student Scenario
by: Wang, Hui, et al.
Published: (2024)
by: Wang, Hui, et al.
Published: (2024)
A solvable model of learning generative diffusion: theory and insights
by: Cui, Hugo, et al.
Published: (2025)
by: Cui, Hugo, et al.
Published: (2025)
Stochasticity-induced non-Hermitian skin criticality
by: Cheng, Xiaoyu, et al.
Published: (2025)
by: Cheng, Xiaoyu, et al.
Published: (2025)
Machine learning for sustainable geoenergy: uncertainty, physics and decision-ready inference
by: Menke, Hannah P., et al.
Published: (2026)
by: Menke, Hannah P., et al.
Published: (2026)
Stochastic synaptic dynamics under learning
by: Stubenrauch, Jakob, et al.
Published: (2025)
by: Stubenrauch, Jakob, et al.
Published: (2025)
Asymptotics of feature learning in two-layer networks after one gradient-step
by: Cui, Hugo, et al.
Published: (2024)
by: Cui, Hugo, et al.
Published: (2024)
Adaptive kernel predictors from feature-learning infinite limits of neural networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Optimal thresholds and algorithms for a model of multi-modal learning in high dimensions
by: Keup, Christian, et al.
Published: (2024)
by: Keup, Christian, et al.
Published: (2024)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Is Grokking a Computational Glass Relaxation?
by: Zhang, Xiaotian, et al.
Published: (2025)
by: Zhang, Xiaotian, et al.
Published: (2025)
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
by: Albergo, Michael S., et al.
Published: (2023)
by: Albergo, Michael S., et al.
Published: (2023)
A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
by: Mendes, Vicente Conde, et al.
Published: (2026)
by: Mendes, Vicente Conde, et al.
Published: (2026)
On the relationship between equilibria and dynamics in large, random neuronal networks
by: Yang, Xiaoyu, et al.
Published: (2025)
by: Yang, Xiaoyu, et al.
Published: (2025)
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
by: Cowsik, Aditya, et al.
Published: (2024)
by: Cowsik, Aditya, et al.
Published: (2024)
Interpretation of Crystal Energy Landscapes with Kolmogorov-Arnold Networks
by: Zu, Gen, et al.
Published: (2026)
by: Zu, Gen, et al.
Published: (2026)
Liquid and solid layers in a thermal deep learning machine
by: Huang, Gang, et al.
Published: (2025)
by: Huang, Gang, et al.
Published: (2025)
Similar Items
-
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023) -
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
by: Kühn, Marcel, et al.
Published: (2023) -
Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems
by: Angelini, Maria Chiara, et al.
Published: (2023) -
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025) -
Random Matrix Theory for Stochastic Gradient Descent
by: Park, Chanju, et al.
Published: (2024)