Robust low-rank training via approximate orthonormal constraints

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Savostianova, Dayana, Zangrando, Emanuele, Ceruti, Gianluca, Tudisco, Francesco
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915809923694592
author Savostianova, Dayana
Zangrando, Emanuele
Ceruti, Gianluca
Tudisco, Francesco
author_facet Savostianova, Dayana
Zangrando, Emanuele
Ceruti, Gianluca
Tudisco, Francesco
contents With the growth of model and data sizes, a broad effort has been made to design pruning techniques that reduce the resource demand of deep learning pipelines, while retaining model performance. In order to reduce both inference and training costs, a prominent line of work uses low-rank matrix factorizations to represent the network weights. Although able to retain accuracy, we observe that low-rank methods tend to compromise model robustness against adversarial perturbations. By modeling robustness in terms of the condition number of the neural network, we argue that this loss of robustness is due to the exploding singular values of the low-rank weight matrices. Thus, we introduce a robust low-rank training algorithm that maintains the network's weights on the low-rank matrix manifold while simultaneously enforcing approximate orthonormal constraints. The resulting model reduces both training and inference costs while ensuring well-conditioning and thus better adversarial robustness, without compromising model accuracy. This is shown by extensive numerical evidence and by our main approximation theorem that shows the computed robust low-rank network well-approximates the ideal full model, provided a highly performing low-rank sub-network exists.
format Preprint
id arxiv_https___arxiv_org_abs_2306_01485
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Robust low-rank training via approximate orthonormal constraints
Savostianova, Dayana
Zangrando, Emanuele
Ceruti, Gianluca
Tudisco, Francesco
Machine Learning
Artificial Intelligence
Numerical Analysis
With the growth of model and data sizes, a broad effort has been made to design pruning techniques that reduce the resource demand of deep learning pipelines, while retaining model performance. In order to reduce both inference and training costs, a prominent line of work uses low-rank matrix factorizations to represent the network weights. Although able to retain accuracy, we observe that low-rank methods tend to compromise model robustness against adversarial perturbations. By modeling robustness in terms of the condition number of the neural network, we argue that this loss of robustness is due to the exploding singular values of the low-rank weight matrices. Thus, we introduce a robust low-rank training algorithm that maintains the network's weights on the low-rank matrix manifold while simultaneously enforcing approximate orthonormal constraints. The resulting model reduces both training and inference costs while ensuring well-conditioning and thus better adversarial robustness, without compromising model accuracy. This is shown by extensive numerical evidence and by our main approximation theorem that shows the computed robust low-rank network well-approximates the ideal full model, provided a highly performing low-rank sub-network exists.
title Robust low-rank training via approximate orthonormal constraints
topic Machine Learning
Artificial Intelligence
Numerical Analysis
url https://arxiv.org/abs/2306.01485