Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liang, Shuang, Montúfar, Guido
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912886393143296
author Liang, Shuang
Montúfar, Guido
author_facet Liang, Shuang
Montúfar, Guido
contents We examine gradient descent in matrix factorization and show that under large step sizes the parameter space develops a fractal structure. We derive the exact critical step size for convergence in scalar-vector factorization and show that near criticality the selected minimizer depends sensitively on the initialization. Moreover, we show that adding regularization amplifies this sensitivity, generating a fractal boundary between initializations that converge and those that diverge. The analysis extends to general matrix factorization with orthogonal initialization. Our findings reveal that near-critical step sizes induce a chaotic regime of gradient descent where the training outcome is unpredictable and there are no simple implicit biases, such as towards balancedness, minimum norm, or flatness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
Liang, Shuang
Montúfar, Guido
Machine Learning
We examine gradient descent in matrix factorization and show that under large step sizes the parameter space develops a fractal structure. We derive the exact critical step size for convergence in scalar-vector factorization and show that near criticality the selected minimizer depends sensitively on the initialization. Moreover, we show that adding regularization amplifies this sensitivity, generating a fractal boundary between initializations that converge and those that diverge. The analysis extends to general matrix factorization with orthogonal initialization. Our findings reveal that near-critical step sizes induce a chaotic regime of gradient descent where the training outcome is unpredictable and there are no simple implicit biases, such as towards balancedness, minimum norm, or flatness.
title Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
topic Machine Learning
url https://arxiv.org/abs/2509.25351