Online Laplace Model Selection Revisited

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lin, Jihao Andreas, Antorán, Javier, Hernández-Lobato, José Miguel
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913189823774720
author Lin, Jihao Andreas
Antorán, Javier
Hernández-Lobato, José Miguel
author_facet Lin, Jihao Andreas
Antorán, Javier
Hernández-Lobato, José Miguel
contents The Laplace approximation provides a closed-form model selection objective for neural networks (NN). Online variants, which optimise NN parameters jointly with hyperparameters, like weight decay strength, have seen renewed interest in the Bayesian deep learning community. However, these methods violate Laplace's method's critical assumption that the approximation is performed around a mode of the loss, calling into question their soundness. This work re-derives online Laplace methods, showing them to target a variational bound on a mode-corrected variant of the Laplace evidence which does not make stationarity assumptions. Online Laplace and its mode-corrected counterpart share stationary points where 1. the NN parameters are a maximum a posteriori, satisfying the Laplace method's assumption, and 2. the hyperparameters maximise the Laplace evidence, motivating online methods. We demonstrate that these optima are roughly attained in practise by online algorithms using full-batch gradient descent on UCI regression datasets. The optimised hyperparameters prevent overfitting and outperform validation-based early stopping.
format Preprint
id arxiv_https___arxiv_org_abs_2307_06093
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Online Laplace Model Selection Revisited
Lin, Jihao Andreas
Antorán, Javier
Hernández-Lobato, José Miguel
Machine Learning
The Laplace approximation provides a closed-form model selection objective for neural networks (NN). Online variants, which optimise NN parameters jointly with hyperparameters, like weight decay strength, have seen renewed interest in the Bayesian deep learning community. However, these methods violate Laplace's method's critical assumption that the approximation is performed around a mode of the loss, calling into question their soundness. This work re-derives online Laplace methods, showing them to target a variational bound on a mode-corrected variant of the Laplace evidence which does not make stationarity assumptions. Online Laplace and its mode-corrected counterpart share stationary points where 1. the NN parameters are a maximum a posteriori, satisfying the Laplace method's assumption, and 2. the hyperparameters maximise the Laplace evidence, motivating online methods. We demonstrate that these optima are roughly attained in practise by online algorithms using full-batch gradient descent on UCI regression datasets. The optimised hyperparameters prevent overfitting and outperform validation-based early stopping.
title Online Laplace Model Selection Revisited
topic Machine Learning
url https://arxiv.org/abs/2307.06093