Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Dibyakanti, Jha, Samyak, Mukherjee, Anirbit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913997141311488
author Kumar, Dibyakanti
Jha, Samyak
Mukherjee, Anirbit
author_facet Kumar, Dibyakanti
Jha, Samyak
Mukherjee, Anirbit
contents In this work, we will establish that the Langevin Monte-Carlo algorithm can learn depth-2 neural nets of any size and for any data and we give non-asymptotic convergence rates for it. We achieve this via showing that in q-Renyi divergence, the iterates of Langevin Monte Carlo converge to the Gibbs distribution of Frobenius norm regularized losses for any of these nets, when using smooth activations and in both classification and regression settings. Most critically, the amount of regularization needed for our results is independent of the size of the net. This result achieves a synthesis of several recent observations about isoperimetry conditions under which LMC converges and that two-layer neural loss functions can always be regularized by a certain constant amount such that they satisfy the Villani conditions, and thus their Gibbs measures satisfy a Poincare inequality.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10428
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
Kumar, Dibyakanti
Jha, Samyak
Mukherjee, Anirbit
Machine Learning
Functional Analysis
Probability
In this work, we will establish that the Langevin Monte-Carlo algorithm can learn depth-2 neural nets of any size and for any data and we give non-asymptotic convergence rates for it. We achieve this via showing that in q-Renyi divergence, the iterates of Langevin Monte Carlo converge to the Gibbs distribution of Frobenius norm regularized losses for any of these nets, when using smooth activations and in both classification and regression settings. Most critically, the amount of regularization needed for our results is independent of the size of the net. This result achieves a synthesis of several recent observations about isoperimetry conditions under which LMC converges and that two-layer neural loss functions can always be regularized by a certain constant amount such that they satisfy the Villani conditions, and thus their Gibbs measures satisfy a Poincare inequality.
title Langevin Monte-Carlo Provably Learns Depth Two Neural Nets at Any Size and Data
topic Machine Learning
Functional Analysis
Probability
url https://arxiv.org/abs/2503.10428