Optimal Asymptotic Rates for (Stochastic) Gradient Descent under the Local PL-Condition: A Geometric Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kassing, Sebastian, Kruse, Thomas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911685595365376
author Kassing, Sebastian
Kruse, Thomas
author_facet Kassing, Sebastian
Kruse, Thomas
contents Stochastic gradient descent (SGD) has been studied extensively over the past decades due to its simplicity and broad applicability in machine learning. In this work, we analyze the local behavior of gradient descent and stochastic gradient descent for minimizing $C^2$-functions that satisfy the Polyak-Lojasiewicz (PL) inequality and under a multiplicative gradient noise model motivated by overparameterized neural networks. Using a geometric interpretation of the PL-condition, we prove a simple yet surprising fact: in this possibly non-convex setting, the asymptotic convergence rate of (S)GD matches the rate obtained for strongly convex quadratics.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14663
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimal Asymptotic Rates for (Stochastic) Gradient Descent under the Local PL-Condition: A Geometric Approach
Kassing, Sebastian
Kruse, Thomas
Optimization and Control
Probability
Machine Learning
Primary 90C26, Secondary 90C15, 62L20, 90C30
Stochastic gradient descent (SGD) has been studied extensively over the past decades due to its simplicity and broad applicability in machine learning. In this work, we analyze the local behavior of gradient descent and stochastic gradient descent for minimizing $C^2$-functions that satisfy the Polyak-Lojasiewicz (PL) inequality and under a multiplicative gradient noise model motivated by overparameterized neural networks. Using a geometric interpretation of the PL-condition, we prove a simple yet surprising fact: in this possibly non-convex setting, the asymptotic convergence rate of (S)GD matches the rate obtained for strongly convex quadratics.
title Optimal Asymptotic Rates for (Stochastic) Gradient Descent under the Local PL-Condition: A Geometric Approach
topic Optimization and Control
Probability
Machine Learning
Primary 90C26, Secondary 90C15, 62L20, 90C30
url https://arxiv.org/abs/2605.14663