Phase diagram and eigenvalue dynamics of stochastic gradient descent in multilayer neural networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Park, Chanju, Lucini, Biagio, Aarts, Gert
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911313552211968
author Park, Chanju
Lucini, Biagio
Aarts, Gert
author_facet Park, Chanju
Lucini, Biagio
Aarts, Gert
contents Hyperparameter tuning is one of the essential steps to guarantee the convergence of machine learning models. We argue that intuition about the optimal choice of hyperparameters for stochastic gradient descent can be obtained by studying a neural network's phase diagram, in which each phase is characterised by distinctive dynamics of the singular values of weight matrices. Taking inspiration from disordered systems, we start from the observation that the loss landscape of a multilayer neural network with mean squared error can be interpreted as a disordered system in feature space, where the learnt features are mapped to soft spin degrees of freedom, the initial variance of the weight matrices is interpreted as the strength of the disorder, and temperature is given by the ratio of the learning rate and the batch size. As the model is trained, three phases can be identified, in which the dynamics of weight matrices is qualitatively different. Employing a Langevin equation for stochastic gradient descent, previously derived using Dyson Brownian motion, we demonstrate that the three dynamical regimes can be classified effectively, providing practical guidance for the choice of hyperparameters of the optimiser.
format Preprint
id arxiv_https___arxiv_org_abs_2509_01349
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Phase diagram and eigenvalue dynamics of stochastic gradient descent in multilayer neural networks
Park, Chanju
Lucini, Biagio
Aarts, Gert
Disordered Systems and Neural Networks
Machine Learning
High Energy Physics - Lattice
Hyperparameter tuning is one of the essential steps to guarantee the convergence of machine learning models. We argue that intuition about the optimal choice of hyperparameters for stochastic gradient descent can be obtained by studying a neural network's phase diagram, in which each phase is characterised by distinctive dynamics of the singular values of weight matrices. Taking inspiration from disordered systems, we start from the observation that the loss landscape of a multilayer neural network with mean squared error can be interpreted as a disordered system in feature space, where the learnt features are mapped to soft spin degrees of freedom, the initial variance of the weight matrices is interpreted as the strength of the disorder, and temperature is given by the ratio of the learning rate and the batch size. As the model is trained, three phases can be identified, in which the dynamics of weight matrices is qualitatively different. Employing a Langevin equation for stochastic gradient descent, previously derived using Dyson Brownian motion, we demonstrate that the three dynamical regimes can be classified effectively, providing practical guidance for the choice of hyperparameters of the optimiser.
title Phase diagram and eigenvalue dynamics of stochastic gradient descent in multilayer neural networks
topic Disordered Systems and Neural Networks
Machine Learning
High Energy Physics - Lattice
url https://arxiv.org/abs/2509.01349