Global Convergence of Four-Layer Matrix Factorization under Random Initialization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luo, Minrui, Xu, Weihang, Gao, Xiang, Fazel, Maryam, Du, Simon Shaolei
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918208828604416
author Luo, Minrui
Xu, Weihang
Gao, Xiang
Fazel, Maryam
Du, Simon Shaolei
author_facet Luo, Minrui
Xu, Weihang
Gao, Xiang
Fazel, Maryam
Du, Simon Shaolei
contents Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is well-established, no global convergence guarantee for general deep matrix factorization under random initialization has been established to date. To address this gap, we provide a polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization, given certain conditions on the target matrix and a standard balanced regularization term. Our analysis employs new techniques to show saddle-avoidance properties of gradient decent dynamics, and extends previous theories to characterize the change in eigenvalues of layer weights.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Global Convergence of Four-Layer Matrix Factorization under Random Initialization
Luo, Minrui
Xu, Weihang
Gao, Xiang
Fazel, Maryam
Du, Simon Shaolei
Optimization and Control
Machine Learning
Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is well-established, no global convergence guarantee for general deep matrix factorization under random initialization has been established to date. To address this gap, we provide a polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization, given certain conditions on the target matrix and a standard balanced regularization term. Our analysis employs new techniques to show saddle-avoidance properties of gradient decent dynamics, and extends previous theories to characterize the change in eigenvalues of layer weights.
title Global Convergence of Four-Layer Matrix Factorization under Random Initialization
topic Optimization and Control
Machine Learning
url https://arxiv.org/abs/2511.09925