On the different regimes of Stochastic Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Sclocchi, Antonio, Wyart, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024)
by: Sclocchi, Antonio, et al.
Published: (2024)
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024)
by: Tomasini, Umberto, et al.
Published: (2024)
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)
by: Cagnetta, Francesco, et al.
Published: (2025)
Towards a theory of how the structure of language is acquired by deep neural networks
by: Cagnetta, Francesco, et al.
Published: (2024)
by: Cagnetta, Francesco, et al.
Published: (2024)
Deep networks learn to parse uniform-depth context-free languages from local statistics
by: Parley, Jack T., et al.
Published: (2026)
by: Parley, Jack T., et al.
Published: (2026)
Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
by: Kühn, Marcel, et al.
Published: (2023)
by: Kühn, Marcel, et al.
Published: (2023)
Sampling Data with Chains of Forward-Backward Diffusion Steps
by: Kang, Hyunmo, et al.
Published: (2026)
by: Kang, Hyunmo, et al.
Published: (2026)
On the Emergence of Linear Analogies in Word Embeddings
by: Korchinski, Daniel J., et al.
Published: (2025)
by: Korchinski, Daniel J., et al.
Published: (2025)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Convergence Acceleration of Markov Chain Monte Carlo-based Gradient Descent by Deep Unfolding
by: Hagiwara, Ryo, et al.
Published: (2024)
by: Hagiwara, Ryo, et al.
Published: (2024)
Symmetry in language statistics shapes the geometry of model representations
by: Karkada, Dhruva, et al.
Published: (2026)
by: Karkada, Dhruva, et al.
Published: (2026)
Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems
by: Angelini, Maria Chiara, et al.
Published: (2023)
by: Angelini, Maria Chiara, et al.
Published: (2023)
Random Matrix Theory for Stochastic Gradient Descent
by: Park, Chanju, et al.
Published: (2024)
by: Park, Chanju, et al.
Published: (2024)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
by: Veiga, Rodrigo, et al.
Published: (2024)
by: Veiga, Rodrigo, et al.
Published: (2024)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Quantum Equilibrium Propagation: Gradient-Descent Training of Quantum Systems
by: Scellier, Benjamin
Published: (2024)
by: Scellier, Benjamin
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Short-range depinning in the presence of velocity-weakening
by: de Geus, Tom W. J., et al.
Published: (2024)
by: de Geus, Tom W. J., et al.
Published: (2024)
Dynamical heterogeneities of thermal creep in pinned interfaces
by: de Geus, Tom W. J., et al.
Published: (2024)
by: de Geus, Tom W. J., et al.
Published: (2024)
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
by: Albergo, Michael S., et al.
Published: (2023)
by: Albergo, Michael S., et al.
Published: (2023)
Microscopic description of the intermittent dynamics driving logarithmic creep
by: Korchinski, Daniel J., et al.
Published: (2024)
by: Korchinski, Daniel J., et al.
Published: (2024)
Fundamental operating regimes, hyper-parameter fine-tuning and glassiness: towards an interpretable replica-theory for trained restricted Boltzmann machines
by: Fachechi, Alberto, et al.
Published: (2024)
by: Fachechi, Alberto, et al.
Published: (2024)
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
by: Martin, Simon, et al.
Published: (2026)
by: Martin, Simon, et al.
Published: (2026)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023)
by: Patel, Nishil, et al.
Published: (2023)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
How does training shape the Riemannian geometry of neural network representations?
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
Grokking as a First Order Phase Transition in Two Layer Networks
by: Rubin, Noa, et al.
Published: (2023)
by: Rubin, Noa, et al.
Published: (2023)
A universal approximation theorem for nonlinear resistive networks
by: Scellier, Benjamin, et al.
Published: (2023)
by: Scellier, Benjamin, et al.
Published: (2023)
High-dimensional Asymptotics of Denoising Autoencoders
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Training neural networks with structured noise improves classification and generalization
by: Benedetti, Marco, et al.
Published: (2023)
by: Benedetti, Marco, et al.
Published: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Deep neural networks from the perspective of ergodic theory
by: Zhang, Fan
Published: (2023)
by: Zhang, Fan
Published: (2023)
Field theory for optimal signal propagation in ResNets
by: Fischer, Kirsten, et al.
Published: (2023)
by: Fischer, Kirsten, et al.
Published: (2023)
DCEM: A deep complementary energy method for solid mechanics
by: Wang, Yizheng, et al.
Published: (2023)
by: Wang, Yizheng, et al.
Published: (2023)
Similar Items
-
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
by: Cagnetta, Francesco, et al.
Published: (2025) -
Probing the Latent Hierarchical Structure of Data via Diffusion Models
by: Sclocchi, Antonio, et al.
Published: (2024) -
A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Data
by: Sclocchi, Antonio, et al.
Published: (2024) -
How Deep Networks Learn Sparse and Hierarchical Data: the Sparse Random Hierarchy Model
by: Tomasini, Umberto, et al.
Published: (2024) -
Learning curves theory for hierarchically compositional data with power-law distributed features
by: Cagnetta, Francesco, et al.
Published: (2025)