Variational Deep Learning via Implicit Regularization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wenger, Jonathan, Coker, Beau, Marusic, Juraj, Cunningham, John P.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917339582169088
author Wenger, Jonathan
Coker, Beau
Marusic, Juraj
Cunningham, John P.
author_facet Wenger, Jonathan
Coker, Beau
Marusic, Juraj
Cunningham, John P.
contents Modern deep learning models generalize remarkably well in-distribution, despite being overparametrized and trained with little to no explicit regularization. Instead, current theory credits implicit regularization imposed by the choice of architecture, hyperparameters, and optimization procedure. However, deep neural networks can be surprisingly non-robust, resulting in overconfident predictions and poor out-of-distribution generalization. Bayesian deep learning addresses this via model averaging, but typically requires significant computational resources as well as carefully elicited priors to avoid overriding the benefits of implicit regularization. Instead, in this work, we propose to regularize variational neural networks solely by relying on the implicit bias of (stochastic) gradient descent. We theoretically characterize this inductive bias in overparametrized linear models as generalized variational inference and demonstrate the importance of the choice of parametrization. Empirically, our approach demonstrates strong in- and out-of-distribution performance without additional hyperparameter tuning and with minimal computational overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20235
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Variational Deep Learning via Implicit Regularization
Wenger, Jonathan
Coker, Beau
Marusic, Juraj
Cunningham, John P.
Machine Learning
Artificial Intelligence
Modern deep learning models generalize remarkably well in-distribution, despite being overparametrized and trained with little to no explicit regularization. Instead, current theory credits implicit regularization imposed by the choice of architecture, hyperparameters, and optimization procedure. However, deep neural networks can be surprisingly non-robust, resulting in overconfident predictions and poor out-of-distribution generalization. Bayesian deep learning addresses this via model averaging, but typically requires significant computational resources as well as carefully elicited priors to avoid overriding the benefits of implicit regularization. Instead, in this work, we propose to regularize variational neural networks solely by relying on the implicit bias of (stochastic) gradient descent. We theoretically characterize this inductive bias in overparametrized linear models as generalized variational inference and demonstrate the importance of the choice of parametrization. Empirically, our approach demonstrates strong in- and out-of-distribution performance without additional hyperparameter tuning and with minimal computational overhead.
title Variational Deep Learning via Implicit Regularization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.20235