Saved in:
Bibliographic Details
Main Authors: Corlouer, Guillaume, Semler, Avi, Strang, Alexander, Oldenziel, Alexander Gietelink
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2604.06366
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908945571905536
author Corlouer, Guillaume
Semler, Avi
Strang, Alexander
Oldenziel, Alexander Gietelink
author_facet Corlouer, Guillaume
Semler, Avi
Strang, Alexander
Oldenziel, Alexander Gietelink
contents Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stochastic gradient descent (SGD) noise on this regime remains poorly understood. We investigate the dynamics of SGD during training of DLNs in the saddle-to-saddle regime. We model the training dynamics as stochastic Langevin dynamics with anisotropic, state-dependent noise. Under the assumption of aligned and balanced weights, we derive an exact decomposition of the dynamics into a system of one-dimensional per-mode stochastic differential equations. This establishes that the maximal diffusion along a mode precedes the corresponding feature being completely learned. We also derive the stationary distribution of SGD for each mode: in the absence of label noise, its marginal distribution along specific features coincides with the stationary distribution of gradient flow, while in the presence of label noise it approximates a Boltzmann distribution. Finally, we confirm experimentally that the theoretical results hold qualitatively even without aligned or balanced weights. These results establish that SGD noise encodes information about the progression of feature learning but does not fundamentally alter the saddle-to-saddle dynamics.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
Corlouer, Guillaume
Semler, Avi
Strang, Alexander
Oldenziel, Alexander Gietelink
Machine Learning
Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stochastic gradient descent (SGD) noise on this regime remains poorly understood. We investigate the dynamics of SGD during training of DLNs in the saddle-to-saddle regime. We model the training dynamics as stochastic Langevin dynamics with anisotropic, state-dependent noise. Under the assumption of aligned and balanced weights, we derive an exact decomposition of the dynamics into a system of one-dimensional per-mode stochastic differential equations. This establishes that the maximal diffusion along a mode precedes the corresponding feature being completely learned. We also derive the stationary distribution of SGD for each mode: in the absence of label noise, its marginal distribution along specific features coincides with the stationary distribution of gradient flow, while in the presence of label noise it approximates a Boltzmann distribution. Finally, we confirm experimentally that the theoretical results hold qualitatively even without aligned or balanced weights. These results establish that SGD noise encodes information about the progression of feature learning but does not fundamentally alter the saddle-to-saddle dynamics.
title Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
topic Machine Learning
url https://arxiv.org/abs/2604.06366