Implicit Regularization Towards Rank Minimization in ReLU Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Timor, Nadav, Vardi, Gal, Shamir, Ohad
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909436729098240
author Timor, Nadav
Vardi, Gal
Shamir, Ohad
author_facet Timor, Nadav
Vardi, Gal
Shamir, Ohad
contents We study the conjectured relationship between the implicit regularization in neural networks, trained with gradient-based methods, and rank minimization of their weight matrices. Previously, it was proved that for linear networks (of depth 2 and vector-valued outputs), gradient flow (GF) w.r.t. the square loss acts as a rank minimization heuristic. However, understanding to what extent this generalizes to nonlinear networks is an open problem. In this paper, we focus on nonlinear ReLU networks, providing several new positive and negative results. On the negative side, we prove (and demonstrate empirically) that, unlike the linear case, GF on ReLU networks may no longer tend to minimize ranks, in a rather strong sense (even approximately, for "most" datasets of size 2). On the positive side, we reveal that ReLU networks of sufficient depth are provably biased towards low-rank solutions in several reasonable settings.
format Preprint
id arxiv_https___arxiv_org_abs_2201_12760
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Implicit Regularization Towards Rank Minimization in ReLU Networks
Timor, Nadav
Vardi, Gal
Shamir, Ohad
Machine Learning
We study the conjectured relationship between the implicit regularization in neural networks, trained with gradient-based methods, and rank minimization of their weight matrices. Previously, it was proved that for linear networks (of depth 2 and vector-valued outputs), gradient flow (GF) w.r.t. the square loss acts as a rank minimization heuristic. However, understanding to what extent this generalizes to nonlinear networks is an open problem. In this paper, we focus on nonlinear ReLU networks, providing several new positive and negative results. On the negative side, we prove (and demonstrate empirically) that, unlike the linear case, GF on ReLU networks may no longer tend to minimize ranks, in a rather strong sense (even approximately, for "most" datasets of size 2). On the positive side, we reveal that ReLU networks of sufficient depth are provably biased towards low-rank solutions in several reasonable settings.
title Implicit Regularization Towards Rank Minimization in ReLU Networks
topic Machine Learning
url https://arxiv.org/abs/2201.12760