Exact Gauss-Newton Optimization for Training Deep Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Korbit, Mikalai, Adeoye, Adeyemi D., Bemporad, Alberto, Zanon, Mario
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912647215054848
author Korbit, Mikalai
Adeoye, Adeyemi D.
Bemporad, Alberto
Zanon, Mario
author_facet Korbit, Mikalai
Adeoye, Adeyemi D.
Bemporad, Alberto
Zanon, Mario
contents We present Exact Gauss-Newton (EGN), a stochastic second-order optimization algorithm that combines the generalized Gauss-Newton (GN) Hessian approximation with low-rank linear algebra to compute the descent direction. Leveraging the Duncan-Guttman matrix identity, the parameter update is obtained by factorizing a matrix which has the size of the mini-batch. This is particularly advantageous for large-scale machine learning problems where the dimension of the neural network parameter vector is several orders of magnitude larger than the batch size. Additionally, we show how improvements such as line search, adaptive regularization, and momentum can be seamlessly added to EGN to further accelerate the algorithm. Moreover, under mild assumptions, we prove that our algorithm converges in expectation to a stationary point of the objective. Finally, our numerical experiments demonstrate that EGN consistently exceeds, or at most matches the generalization performance of well-tuned SGD, Adam, GAF, SQN, and SGN optimizers across various supervised and reinforcement learning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14402
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exact Gauss-Newton Optimization for Training Deep Neural Networks
Korbit, Mikalai
Adeoye, Adeyemi D.
Bemporad, Alberto
Zanon, Mario
Machine Learning
We present Exact Gauss-Newton (EGN), a stochastic second-order optimization algorithm that combines the generalized Gauss-Newton (GN) Hessian approximation with low-rank linear algebra to compute the descent direction. Leveraging the Duncan-Guttman matrix identity, the parameter update is obtained by factorizing a matrix which has the size of the mini-batch. This is particularly advantageous for large-scale machine learning problems where the dimension of the neural network parameter vector is several orders of magnitude larger than the batch size. Additionally, we show how improvements such as line search, adaptive regularization, and momentum can be seamlessly added to EGN to further accelerate the algorithm. Moreover, under mild assumptions, we prove that our algorithm converges in expectation to a stationary point of the objective. Finally, our numerical experiments demonstrate that EGN consistently exceeds, or at most matches the generalization performance of well-tuned SGD, Adam, GAF, SQN, and SGN optimizers across various supervised and reinforcement learning tasks.
title Exact Gauss-Newton Optimization for Training Deep Neural Networks
topic Machine Learning
url https://arxiv.org/abs/2405.14402