Debiasing Mini-Batch Quadratics for Applications in Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tatzel, Lukas, Mucsányi, Bálint, Hackel, Osane, Hennig, Philipp
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913667402956800
author Tatzel, Lukas
Mucsányi, Bálint
Hackel, Osane
Hennig, Philipp
author_facet Tatzel, Lukas
Mucsányi, Bálint
Hackel, Osane
Hennig, Philipp
contents Quadratic approximations form a fundamental building block of machine learning methods. E.g., second-order optimizers try to find the Newton step into the minimum of a local quadratic proxy to the objective function; and the second-order approximation of a network's loss function can be used to quantify the uncertainty of its outputs via the Laplace approximation. When computations on the entire training set are intractable - typical for deep learning - the relevant quantities are computed on mini-batches. This, however, distorts and biases the shape of the associated stochastic quadratic approximations in an intricate way with detrimental effects on applications. In this paper, we (i) show that this bias introduces a systematic error, (ii) provide a theoretical explanation for it, (iii) explain its relevance for second-order optimization and uncertainty quantification via the Laplace approximation in deep learning, and (iv) develop and evaluate debiasing strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14325
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Debiasing Mini-Batch Quadratics for Applications in Deep Learning
Tatzel, Lukas
Mucsányi, Bálint
Hackel, Osane
Hennig, Philipp
Machine Learning
Quadratic approximations form a fundamental building block of machine learning methods. E.g., second-order optimizers try to find the Newton step into the minimum of a local quadratic proxy to the objective function; and the second-order approximation of a network's loss function can be used to quantify the uncertainty of its outputs via the Laplace approximation. When computations on the entire training set are intractable - typical for deep learning - the relevant quantities are computed on mini-batches. This, however, distorts and biases the shape of the associated stochastic quadratic approximations in an intricate way with detrimental effects on applications. In this paper, we (i) show that this bias introduces a systematic error, (ii) provide a theoretical explanation for it, (iii) explain its relevance for second-order optimization and uncertainty quantification via the Laplace approximation in deep learning, and (iv) develop and evaluate debiasing strategies.
title Debiasing Mini-Batch Quadratics for Applications in Deep Learning
topic Machine Learning
url https://arxiv.org/abs/2410.14325