Gradient Span Algorithms Make Predictable Progress in High Dimension

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Benning, Felix, Döring, Leif
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916437359067136
author Benning, Felix
Döring, Leif
author_facet Benning, Felix
Döring, Leif
contents We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. In particular, this result explains the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. The distributional assumption of (non-stationary) isotropic Gaussian random functions we use is sufficiently general to serve as realistic model for machine learning training but also encompass spin glasses and random quadratic functions.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09973
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Gradient Span Algorithms Make Predictable Progress in High Dimension
Benning, Felix
Döring, Leif
Machine Learning
Optimization and Control
Probability
60F99, 68T01, 82D30
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. In particular, this result explains the counterintuitive phenomenon that different training runs of many large machine learning models result in approximately equal cost curves despite random initialization on a complicated non-convex landscape. The distributional assumption of (non-stationary) isotropic Gaussian random functions we use is sufficiently general to serve as realistic model for machine learning training but also encompass spin glasses and random quadratic functions.
title Gradient Span Algorithms Make Predictable Progress in High Dimension
topic Machine Learning
Optimization and Control
Probability
60F99, 68T01, 82D30
url https://arxiv.org/abs/2410.09973