Saved in:
Bibliographic Details
Main Authors: Nenov, Rossen, Haider, Daniel, Balazs, Peter
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.00169
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910626623782912
author Nenov, Rossen
Haider, Daniel
Balazs, Peter
author_facet Nenov, Rossen
Haider, Daniel
Balazs, Peter
contents Maintaining numerical stability in machine learning models is crucial for their reliability and performance. One approach to maintain stability of a network layer is to integrate the condition number of the weight matrix as a regularizing term into the optimization algorithm. However, due to its discontinuous nature and lack of differentiability the condition number is not suitable for a gradient descent approach. This paper introduces a novel regularizer that is provably differentiable almost everywhere and promotes matrices with low condition numbers. In particular, we derive a formula for the gradient of this regularizer which can be easily implemented and integrated into existing optimization algorithms. We show the advantages of this approach for noisy classification and denoising of MNIST images.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00169
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle (Almost) Smooth Sailing: Towards Numerical Stability of Neural Networks Through Differentiable Regularization of the Condition Number
Nenov, Rossen
Haider, Daniel
Balazs, Peter
Machine Learning
Optimization and Control
Maintaining numerical stability in machine learning models is crucial for their reliability and performance. One approach to maintain stability of a network layer is to integrate the condition number of the weight matrix as a regularizing term into the optimization algorithm. However, due to its discontinuous nature and lack of differentiability the condition number is not suitable for a gradient descent approach. This paper introduces a novel regularizer that is provably differentiable almost everywhere and promotes matrices with low condition numbers. In particular, we derive a formula for the gradient of this regularizer which can be easily implemented and integrated into existing optimization algorithms. We show the advantages of this approach for noisy classification and denoising of MNIST images.
title (Almost) Smooth Sailing: Towards Numerical Stability of Neural Networks Through Differentiable Regularization of the Condition Number
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2410.00169