Low-Rank Filtering and Smoothing for Sequential Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sliwa, Joanna, Schneider, Frank, Bosch, Nathanael, Kristiadi, Agustinus, Hennig, Philipp
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914210801254400
author Sliwa, Joanna
Schneider, Frank
Bosch, Nathanael
Kristiadi, Agustinus
Hennig, Philipp
author_facet Sliwa, Joanna
Schneider, Frank
Bosch, Nathanael
Kristiadi, Agustinus
Hennig, Philipp
contents Learning multiple tasks sequentially requires neural networks to balance retaining knowledge, yet being flexible enough to adapt to new tasks. Regularizing network parameters is a common approach, but it rarely incorporates prior knowledge about task relationships, and limits information flow to future tasks only. We propose a Bayesian framework that treats the network's parameters as the state space of a nonlinear Gaussian model, unlocking two key capabilities: (1) A principled way to encode domain knowledge about task relationships, allowing, e.g., control over which layers should adapt between tasks. (2) A novel application of Bayesian smoothing, allowing task-specific models to also incorporate knowledge from models learned later. This does not require direct access to their data, which is crucial, e.g., for privacy-critical applications. These capabilities rely on efficient filtering and smoothing operations, for which we propose diagonal plus low-rank approximations of the precision matrix in the Laplace approximation (LR-LGF). Empirical results demonstrate the efficiency of LR-LGF and the benefits of the unlocked capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06800
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-Rank Filtering and Smoothing for Sequential Deep Learning
Sliwa, Joanna
Schneider, Frank
Bosch, Nathanael
Kristiadi, Agustinus
Hennig, Philipp
Machine Learning
Learning multiple tasks sequentially requires neural networks to balance retaining knowledge, yet being flexible enough to adapt to new tasks. Regularizing network parameters is a common approach, but it rarely incorporates prior knowledge about task relationships, and limits information flow to future tasks only. We propose a Bayesian framework that treats the network's parameters as the state space of a nonlinear Gaussian model, unlocking two key capabilities: (1) A principled way to encode domain knowledge about task relationships, allowing, e.g., control over which layers should adapt between tasks. (2) A novel application of Bayesian smoothing, allowing task-specific models to also incorporate knowledge from models learned later. This does not require direct access to their data, which is crucial, e.g., for privacy-critical applications. These capabilities rely on efficient filtering and smoothing operations, for which we propose diagonal plus low-rank approximations of the precision matrix in the Laplace approximation (LR-LGF). Empirical results demonstrate the efficiency of LR-LGF and the benefits of the unlocked capabilities.
title Low-Rank Filtering and Smoothing for Sequential Deep Learning
topic Machine Learning
url https://arxiv.org/abs/2410.06800