Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Karnik, Santhosh, Veselovska, Anna, Iwen, Mark, Krahmer, Felix
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913558243049472
author Karnik, Santhosh
Veselovska, Anna
Iwen, Mark
Krahmer, Felix
author_facet Karnik, Santhosh
Veselovska, Anna
Iwen, Mark
Krahmer, Felix
contents We provide a rigorous analysis of implicit regularization in an overparametrized tensor factorization problem beyond the lazy training regime. For matrix factorization problems, this phenomenon has been studied in a number of works. A particular challenge has been to design universal initialization strategies which provably lead to implicit regularization in gradient-descent methods. At the same time, it has been argued by Cohen et. al. 2016 that more general classes of neural networks can be captured by considering tensor factorizations. However, in the tensor case, implicit regularization has only been rigorously established for gradient flow or in the lazy training regime. In this paper, we prove the first tensor result of its kind for gradient descent rather than gradient flow. We focus on the tubal tensor product and the associated notion of low tubal rank, encouraged by the relevance of this model for image data. We establish that gradient descent in an overparametrized tensor factorization model with a small random initialization exhibits an implicit bias towards solutions of low tubal rank. Our theoretical findings are illustrated in an extensive set of numerical simulations show-casing the dynamics predicted by our theory as well as the crucial role of using a small random initialization.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16247
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
Karnik, Santhosh
Veselovska, Anna
Iwen, Mark
Krahmer, Felix
Machine Learning
Optimization and Control
Statistics Theory
We provide a rigorous analysis of implicit regularization in an overparametrized tensor factorization problem beyond the lazy training regime. For matrix factorization problems, this phenomenon has been studied in a number of works. A particular challenge has been to design universal initialization strategies which provably lead to implicit regularization in gradient-descent methods. At the same time, it has been argued by Cohen et. al. 2016 that more general classes of neural networks can be captured by considering tensor factorizations. However, in the tensor case, implicit regularization has only been rigorously established for gradient flow or in the lazy training regime. In this paper, we prove the first tensor result of its kind for gradient descent rather than gradient flow. We focus on the tubal tensor product and the associated notion of low tubal rank, encouraged by the relevance of this model for image data. We establish that gradient descent in an overparametrized tensor factorization model with a small random initialization exhibits an implicit bias towards solutions of low tubal rank. Our theoretical findings are illustrated in an extensive set of numerical simulations show-casing the dynamics predicted by our theory as well as the crucial role of using a small random initialization.
title Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
topic Machine Learning
Optimization and Control
Statistics Theory
url https://arxiv.org/abs/2410.16247