Saved in:
Bibliographic Details
Main Authors: Eschenhagen, Runa, Immer, Alexander, Turner, Richard E., Schneider, Frank, Hennig, Philipp
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2311.00636
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910294462169088
author Eschenhagen, Runa
Immer, Alexander
Turner, Richard E.
Schneider, Frank
Hennig, Philipp
author_facet Eschenhagen, Runa
Immer, Alexander
Turner, Richard E.
Schneider, Frank
Hennig, Philipp
contents The core components of many modern neural network architectures, such as transformers, convolutional, or graph neural networks, can be expressed as linear layers with $\textit{weight-sharing}$. Kronecker-Factored Approximate Curvature (K-FAC), a second-order optimisation method, has shown promise to speed up neural network training and thereby reduce computational costs. However, there is currently no framework to apply it to generic architectures, specifically ones with linear weight-sharing layers. In this work, we identify two different settings of linear weight-sharing layers which motivate two flavours of K-FAC -- $\textit{expand}$ and $\textit{reduce}$. We show that they are exact for deep linear networks with weight-sharing in their respective setting. Notably, K-FAC-reduce is generally faster than K-FAC-expand, which we leverage to speed up automatic hyperparameter selection via optimising the marginal likelihood for a Wide ResNet. Finally, we observe little difference between these two K-FAC variations when using them to train both a graph neural network and a vision transformer. However, both variations are able to reach a fixed validation metric target in $50$-$75\%$ of the number of steps of a first-order reference run, which translates into a comparable improvement in wall-clock time. This highlights the potential of applying K-FAC to modern neural network architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2311_00636
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
Eschenhagen, Runa
Immer, Alexander
Turner, Richard E.
Schneider, Frank
Hennig, Philipp
Machine Learning
The core components of many modern neural network architectures, such as transformers, convolutional, or graph neural networks, can be expressed as linear layers with $\textit{weight-sharing}$. Kronecker-Factored Approximate Curvature (K-FAC), a second-order optimisation method, has shown promise to speed up neural network training and thereby reduce computational costs. However, there is currently no framework to apply it to generic architectures, specifically ones with linear weight-sharing layers. In this work, we identify two different settings of linear weight-sharing layers which motivate two flavours of K-FAC -- $\textit{expand}$ and $\textit{reduce}$. We show that they are exact for deep linear networks with weight-sharing in their respective setting. Notably, K-FAC-reduce is generally faster than K-FAC-expand, which we leverage to speed up automatic hyperparameter selection via optimising the marginal likelihood for a Wide ResNet. Finally, we observe little difference between these two K-FAC variations when using them to train both a graph neural network and a vision transformer. However, both variations are able to reach a fixed validation metric target in $50$-$75\%$ of the number of steps of a first-order reference run, which translates into a comparable improvement in wall-clock time. This highlights the potential of applying K-FAC to modern neural network architectures.
title Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
topic Machine Learning
url https://arxiv.org/abs/2311.00636