Gradients of Functions of Large Matrices

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Krämer, Nicholas, Moreno-Muñoz, Pablo, Roy, Hrittik, Hauberg, Søren
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913563235319808
author Krämer, Nicholas
Moreno-Muñoz, Pablo
Roy, Hrittik
Hauberg, Søren
author_facet Krämer, Nicholas
Moreno-Muñoz, Pablo
Roy, Hrittik
Hauberg, Søren
contents Tuning scientific and probabilistic machine learning models $-$ for example, partial differential equations, Gaussian processes, or Bayesian neural networks $-$ often relies on evaluating functions of matrices whose size grows with the data set or the number of parameters. While the state-of-the-art for evaluating these quantities is almost always based on Lanczos and Arnoldi iterations, the present work is the first to explain how to differentiate these workhorses of numerical linear algebra efficiently. To get there, we derive previously unknown adjoint systems for Lanczos and Arnoldi iterations, implement them in JAX, and show that the resulting code can compete with Diffrax when it comes to differentiating PDEs, GPyTorch for selecting Gaussian process models and beats standard factorisation methods for calibrating Bayesian neural networks. All this is achieved without any problem-specific code optimisation. Find the code at https://github.com/pnkraemer/experiments-lanczos-adjoints and install the library with pip install matfree.
format Preprint
id arxiv_https___arxiv_org_abs_2405_17277
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gradients of Functions of Large Matrices
Krämer, Nicholas
Moreno-Muñoz, Pablo
Roy, Hrittik
Hauberg, Søren
Machine Learning
Numerical Analysis
Tuning scientific and probabilistic machine learning models $-$ for example, partial differential equations, Gaussian processes, or Bayesian neural networks $-$ often relies on evaluating functions of matrices whose size grows with the data set or the number of parameters. While the state-of-the-art for evaluating these quantities is almost always based on Lanczos and Arnoldi iterations, the present work is the first to explain how to differentiate these workhorses of numerical linear algebra efficiently. To get there, we derive previously unknown adjoint systems for Lanczos and Arnoldi iterations, implement them in JAX, and show that the resulting code can compete with Diffrax when it comes to differentiating PDEs, GPyTorch for selecting Gaussian process models and beats standard factorisation methods for calibrating Bayesian neural networks. All this is achieved without any problem-specific code optimisation. Find the code at https://github.com/pnkraemer/experiments-lanczos-adjoints and install the library with pip install matfree.
title Gradients of Functions of Large Matrices
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2405.17277