High-dimensional Limit of SGD for Diagonal Linear Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malaxechebarría, Begoña García, Paquette, Courtney, Fazel, Maryam, Drusvyatskiy, Dmitriy
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909052081012736
author Malaxechebarría, Begoña García
Paquette, Courtney
Fazel, Maryam
Drusvyatskiy, Dmitriy
author_facet Malaxechebarría, Begoña García
Paquette, Courtney
Fazel, Maryam
Drusvyatskiy, Dmitriy
contents Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and generalization properties of neural models. In this work, we show that in the high-dimensional regime, stochastic gradient descent on diagonal linear networks is well-approximated by continuous dynamics governed by a stochastic differential equation (SDE), which explicitly decouples the drift from the gradient noise. We further derive a deterministic partial differential equation whose solution propagates the relevant state of the iterates and characterizes the time evolution of a broad class of observable statistics, including the risk, curvature, and other metrics for optimality. Finally, we show that, under a suitable parametrization, the stochastic dynamics are globally well posed and converge exponentially fast to zero risk with high probability, yielding a fully explicit non-asymptotic description of their long-time behavior. Numerical simulations corroborate our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2605_17177
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle High-dimensional Limit of SGD for Diagonal Linear Networks
Malaxechebarría, Begoña García
Paquette, Courtney
Fazel, Maryam
Drusvyatskiy, Dmitriy
Optimization and Control
Machine Learning
Statistics Theory
Understanding the behavior of stochastic gradient methods is a central problem in modern machine learning. Recent work has highlighted diagonal linear networks as a simplified yet expressive setting for analyzing the optimization and generalization properties of neural models. In this work, we show that in the high-dimensional regime, stochastic gradient descent on diagonal linear networks is well-approximated by continuous dynamics governed by a stochastic differential equation (SDE), which explicitly decouples the drift from the gradient noise. We further derive a deterministic partial differential equation whose solution propagates the relevant state of the iterates and characterizes the time evolution of a broad class of observable statistics, including the risk, curvature, and other metrics for optimality. Finally, we show that, under a suitable parametrization, the stochastic dynamics are globally well posed and converge exponentially fast to zero risk with high probability, yielding a fully explicit non-asymptotic description of their long-time behavior. Numerical simulations corroborate our theoretical findings.
title High-dimensional Limit of SGD for Diagonal Linear Networks
topic Optimization and Control
Machine Learning
Statistics Theory
url https://arxiv.org/abs/2605.17177