Saved in:
Bibliographic Details
Main Authors: Grishina, Ekaterina, Gorbunov, Mikhail, Rakhuba, Maxim
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.11859
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910610190499840
author Grishina, Ekaterina
Gorbunov, Mikhail
Rakhuba, Maxim
author_facet Grishina, Ekaterina
Gorbunov, Mikhail
Rakhuba, Maxim
contents Controlling the spectral norm of the Jacobian matrix, which is related to the convolution operation, has been shown to improve generalization, training stability and robustness in CNNs. Existing methods for computing the norm either tend to overestimate it or their performance may deteriorate quickly with increasing the input and kernel sizes. In this paper, we demonstrate that the tensor version of the spectral norm of a four-dimensional convolution kernel, up to a constant factor, serves as an upper bound for the spectral norm of the Jacobian matrix associated with the convolution operation. This new upper bound is independent of the input image resolution, differentiable and can be efficiently calculated during training. Through experiments, we demonstrate how this new bound can be used to improve the performance of convolutional architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2409_11859
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
Grishina, Ekaterina
Gorbunov, Mikhail
Rakhuba, Maxim
Machine Learning
Controlling the spectral norm of the Jacobian matrix, which is related to the convolution operation, has been shown to improve generalization, training stability and robustness in CNNs. Existing methods for computing the norm either tend to overestimate it or their performance may deteriorate quickly with increasing the input and kernel sizes. In this paper, we demonstrate that the tensor version of the spectral norm of a four-dimensional convolution kernel, up to a constant factor, serves as an upper bound for the spectral norm of the Jacobian matrix associated with the convolution operation. This new upper bound is independent of the input image resolution, differentiable and can be efficiently calculated during training. Through experiments, we demonstrate how this new bound can be used to improve the performance of convolutional architectures.
title Tight and Efficient Upper Bound on Spectral Norm of Convolutional Layers
topic Machine Learning
url https://arxiv.org/abs/2409.11859