BALI: Learning Neural Networks via Bayesian Layerwise Inference

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kurle, Richard, Klushyn, Alexej, Herbrich, Ralf
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929596337750016
author Kurle, Richard
Klushyn, Alexej
Herbrich, Ralf
author_facet Kurle, Richard
Klushyn, Alexej
Herbrich, Ralf
contents We introduce a new method for learning Bayesian neural networks, treating them as a stack of multivariate Bayesian linear regression models. The main idea is to infer the layerwise posterior exactly if we know the target outputs of each layer. We define these pseudo-targets as the layer outputs from the forward pass, updated by the backpropagated gradients of the objective function. The resulting layerwise posterior is a matrix-normal distribution with a Kronecker-factorized covariance matrix, which can be efficiently inverted. Our method extends to the stochastic mini-batch setting using an exponential moving average over natural-parameter terms, thus gradually forgetting older data. The method converges in few iterations and performs as well as or better than leading Bayesian neural network methods on various regression, classification, and out-of-distribution detection benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2411_12102
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BALI: Learning Neural Networks via Bayesian Layerwise Inference
Kurle, Richard
Klushyn, Alexej
Herbrich, Ralf
Machine Learning
We introduce a new method for learning Bayesian neural networks, treating them as a stack of multivariate Bayesian linear regression models. The main idea is to infer the layerwise posterior exactly if we know the target outputs of each layer. We define these pseudo-targets as the layer outputs from the forward pass, updated by the backpropagated gradients of the objective function. The resulting layerwise posterior is a matrix-normal distribution with a Kronecker-factorized covariance matrix, which can be efficiently inverted. Our method extends to the stochastic mini-batch setting using an exponential moving average over natural-parameter terms, thus gradually forgetting older data. The method converges in few iterations and performs as well as or better than leading Bayesian neural network methods on various regression, classification, and out-of-distribution detection benchmarks.
title BALI: Learning Neural Networks via Bayesian Layerwise Inference
topic Machine Learning
url https://arxiv.org/abs/2411.12102