Anti-Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances in Flat Directions
Fuente:
arXiv
Saved in:
| Main Authors: | Kühn, Marcel, Rosenow, Bernd |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026)
by: Kühn, Marcel, et al.
Published: (2026)
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023)
by: Sclocchi, Antonio, et al.
Published: (2023)
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023)
by: Staats, Max, et al.
Published: (2023)
Boundary between noise and information applied to filtering neural network weight matrices
by: Staats, Max, et al.
Published: (2022)
by: Staats, Max, et al.
Published: (2022)
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)
by: Staats, Max, et al.
Published: (2024)
Transient learning dynamics drive escape from sharp valleys in Stochastic Gradient Descent
by: Yang, Ning, et al.
Published: (2026)
by: Yang, Ning, et al.
Published: (2026)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Unified Description of Learning Dynamics in the Soft Committee Machine from Finite to Ultra-Wide Regimes
by: Afanah, Assem, et al.
Published: (2025)
by: Afanah, Assem, et al.
Published: (2025)
Continuous Specialization Transition in the Soft Committee Machine with ReLU Activation
by: Afanah, Assem, et al.
Published: (2026)
by: Afanah, Assem, et al.
Published: (2026)
Convergence Acceleration of Markov Chain Monte Carlo-based Gradient Descent by Deep Unfolding
by: Hagiwara, Ryo, et al.
Published: (2024)
by: Hagiwara, Ryo, et al.
Published: (2024)
Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems
by: Angelini, Maria Chiara, et al.
Published: (2023)
by: Angelini, Maria Chiara, et al.
Published: (2023)
Random Matrix Theory for Stochastic Gradient Descent
by: Park, Chanju, et al.
Published: (2024)
by: Park, Chanju, et al.
Published: (2024)
BBP Phase Transition for an Extensive Number of Outliers
by: Forner, Niklas, et al.
Published: (2025)
by: Forner, Niklas, et al.
Published: (2025)
Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models
by: Atanasov, Alexander, et al.
Published: (2025)
by: Atanasov, Alexander, et al.
Published: (2025)
Stochastic Gradient Flow Dynamics of Test Risk and its Exact Solution for Weak Features
by: Veiga, Rodrigo, et al.
Published: (2024)
by: Veiga, Rodrigo, et al.
Published: (2024)
High-Dimensional Limit of Stochastic Gradient Flow via Dynamical Mean-Field Theory
by: Nishiyama, Sota, et al.
Published: (2026)
by: Nishiyama, Sota, et al.
Published: (2026)
Quantum Equilibrium Propagation: Gradient-Descent Training of Quantum Systems
by: Scellier, Benjamin
Published: (2024)
by: Scellier, Benjamin
Published: (2024)
Analog Physical Systems Can Exhibit Double Descent
by: Dillavou, Sam, et al.
Published: (2025)
by: Dillavou, Sam, et al.
Published: (2025)
Dataset-Free Weight-Initialization on Restricted Boltzmann Machine
by: Yasuda, Muneki, et al.
Published: (2024)
by: Yasuda, Muneki, et al.
Published: (2024)
Soft Quantization: Model Compression Via Weight Coupling
by: Bernstein, Daniel T., et al.
Published: (2026)
by: Bernstein, Daniel T., et al.
Published: (2026)
Learning Stochastic Thermodynamics Directly from Correlation and Trajectory-Fluctuation Currents
by: Lyu, Jinghao, et al.
Published: (2025)
by: Lyu, Jinghao, et al.
Published: (2025)
Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
by: Albergo, Michael S., et al.
Published: (2023)
by: Albergo, Michael S., et al.
Published: (2023)
EB-RANSAC: Random Sample Consensus based on Energy-Based Model
by: Yasuda, Muneki, et al.
Published: (2026)
by: Yasuda, Muneki, et al.
Published: (2026)
Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence
by: Dawid, Anna, et al.
Published: (2023)
by: Dawid, Anna, et al.
Published: (2023)
Stochastic Dynamics of Skyrmions on a Racetrack: Impact of Equilibrium and Nonequilibrium Noise
by: Hlushchenko, Anton V., et al.
Published: (2025)
by: Hlushchenko, Anton V., et al.
Published: (2025)
Noise tolerance via reinforcement: Learning a reinforced quantum dynamics
by: Ramezanpour, Abolfazl
Published: (2025)
by: Ramezanpour, Abolfazl
Published: (2025)
High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks
by: Martin, Simon, et al.
Published: (2026)
by: Martin, Simon, et al.
Published: (2026)
The RL Perceptron: Generalisation Dynamics of Policy Learning in High Dimensions
by: Patel, Nishil, et al.
Published: (2023)
by: Patel, Nishil, et al.
Published: (2023)
The Quantization Model of Neural Scaling
by: Michaud, Eric J., et al.
Published: (2023)
by: Michaud, Eric J., et al.
Published: (2023)
How does training shape the Riemannian geometry of neural network representations?
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
by: Zavatone-Veth, Jacob A., et al.
Published: (2023)
Grokking as a First Order Phase Transition in Two Layer Networks
by: Rubin, Noa, et al.
Published: (2023)
by: Rubin, Noa, et al.
Published: (2023)
A universal approximation theorem for nonlinear resistive networks
by: Scellier, Benjamin, et al.
Published: (2023)
by: Scellier, Benjamin, et al.
Published: (2023)
High-dimensional Asymptotics of Denoising Autoencoders
by: Cui, Hugo, et al.
Published: (2023)
by: Cui, Hugo, et al.
Published: (2023)
Training neural networks with structured noise improves classification and generalization
by: Benedetti, Marco, et al.
Published: (2023)
by: Benedetti, Marco, et al.
Published: (2023)
Grokking as the Transition from Lazy to Rich Training Dynamics
by: Kumar, Tanishq, et al.
Published: (2023)
by: Kumar, Tanishq, et al.
Published: (2023)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Deep neural networks from the perspective of ergodic theory
by: Zhang, Fan
Published: (2023)
by: Zhang, Fan
Published: (2023)
Field theory for optimal signal propagation in ResNets
by: Fischer, Kirsten, et al.
Published: (2023)
by: Fischer, Kirsten, et al.
Published: (2023)
DCEM: A deep complementary energy method for solid mechanics
by: Wang, Yizheng, et al.
Published: (2023)
by: Wang, Yizheng, et al.
Published: (2023)
Applying statistical learning theory to deep learning
by: Gerbelot, Cédric, et al.
Published: (2023)
by: Gerbelot, Cédric, et al.
Published: (2023)
Similar Items
-
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
by: Kühn, Marcel, et al.
Published: (2026) -
On the different regimes of Stochastic Gradient Descent
by: Sclocchi, Antonio, et al.
Published: (2023) -
Enhancing Noise-Robust Losses for Large-Scale Noisy Data Learning
by: Staats, Max, et al.
Published: (2023) -
Boundary between noise and information applied to filtering neural network weight matrices
by: Staats, Max, et al.
Published: (2022) -
Small Singular Values Matter: A Random Matrix Analysis of Transformer Models
by: Staats, Max, et al.
Published: (2024)