Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sadrtdinov, Ildus, Lobacheva, Ekaterina, Klimov, Ivan, Burtsev, Mikhail, Katsnelson, Mikhail I., Vetrov, Dmitry
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909040783654912
author Sadrtdinov, Ildus
Lobacheva, Ekaterina
Klimov, Ivan
Burtsev, Mikhail
Katsnelson, Mikhail I.
Vetrov, Dmitry
author_facet Sadrtdinov, Ildus
Lobacheva, Ekaterina
Klimov, Ivan
Burtsev, Mikhail
Katsnelson, Mikhail I.
Vetrov, Dmitry
contents Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective, we develop a thermodynamic framework to describe the stationary distributions of stochastic gradient descent (SGD) with weight decay for scale-invariant neural networks, a setting that both reflects practical architectures with normalization layers and permits theoretical analysis. We establish analogies between training hyperparameters (e.g., learning rate, weight decay) and thermodynamic variables such as temperature, pressure, and volume. Starting with a simplified isotropic noise model, we uncover a close correspondence between SGD dynamics and ideal gas behavior, validated through theory and simulation. Extending to training of neural networks, we show that key predictions of the framework, including the behavior of stationary entropy, align closely with experimental observations. This framework provides a principled foundation for interpreting training dynamics and may guide future work on hyperparameter tuning and the design of learning rate schedulers.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
Sadrtdinov, Ildus
Lobacheva, Ekaterina
Klimov, Ivan
Burtsev, Mikhail
Katsnelson, Mikhail I.
Vetrov, Dmitry
Machine Learning
Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective, we develop a thermodynamic framework to describe the stationary distributions of stochastic gradient descent (SGD) with weight decay for scale-invariant neural networks, a setting that both reflects practical architectures with normalization layers and permits theoretical analysis. We establish analogies between training hyperparameters (e.g., learning rate, weight decay) and thermodynamic variables such as temperature, pressure, and volume. Starting with a simplified isotropic noise model, we uncover a close correspondence between SGD dynamics and ideal gas behavior, validated through theory and simulation. Extending to training of neural networks, we show that key predictions of the framework, including the behavior of stationary entropy, align closely with experimental observations. This framework provides a principled foundation for interpreting training dynamics and may guide future work on hyperparameter tuning and the design of learning rate schedulers.
title Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?
topic Machine Learning
url https://arxiv.org/abs/2511.07308