Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Oberweis, Noah, Cayci, Semih
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911406492745728
author Oberweis, Noah
Cayci, Semih
author_facet Oberweis, Noah
Cayci, Semih
contents Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence analysis of stochastic gradient Langevin dynamics (SGLD), which is an Itô stochastic differential equation (SDE) approximation of stochastic gradient descent in continuous time, in the lazy training regime. We show that, under regularity conditions on the Hessian of the loss function, SGLD with multiplicative and state-dependent noise (i) yields a non-degenerate kernel throughout the training process with high probability, and (ii) achieves exponential convergence to the empirical risk minimizer in expectation, and we establish finite-time and finite-width bounds on the optimality gap. We corroborate our theoretical findings with numerical examples in the regression setting.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
Oberweis, Noah
Cayci, Semih
Machine Learning
Optimization and Control
Continuous-time models provide important insights into the training dynamics of optimization algorithms in deep learning. In this work, we establish a non-asymptotic convergence analysis of stochastic gradient Langevin dynamics (SGLD), which is an Itô stochastic differential equation (SDE) approximation of stochastic gradient descent in continuous time, in the lazy training regime. We show that, under regularity conditions on the Hessian of the loss function, SGLD with multiplicative and state-dependent noise (i) yields a non-degenerate kernel throughout the training process with high probability, and (ii) achieves exponential convergence to the empirical risk minimizer in expectation, and we establish finite-time and finite-width bounds on the optimality gap. We corroborate our theoretical findings with numerical examples in the regression setting.
title Convergence of Stochastic Gradient Langevin Dynamics in the Lazy Training Regime
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2510.21245