Convergence Analysis of SGD under Expected Smoothness

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kawamoto, Yuta, Iiduka, Hideaki
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909869506822144
author Kawamoto, Yuta
Iiduka, Hideaki
author_facet Kawamoto, Yuta
Iiduka, Hideaki
contents Stochastic gradient descent (SGD) is the workhorse of large-scale learning, yet classical analyses rely on assumptions that can be either too strong (bounded variance) or too coarse (uniform noise). The expected smoothness (ES) condition has emerged as a flexible alternative that ties the second moment of stochastic gradients to the objective value and the full gradient. This paper presents a self-contained convergence analysis of SGD under ES. We (i) refine ES with interpretations and sampling-dependent constants; (ii) derive bounds of the expectation of squared full gradient norm; and (iii) prove $O(1/K)$ rates with explicit residual errors for various step-size schedules. All proofs are given in full detail in the appendix. Our treatment unifies and extends recent threads (Khaled and Richtárik, 2020; Umeda and Iiduka, 2025).
format Preprint
id arxiv_https___arxiv_org_abs_2510_20608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Convergence Analysis of SGD under Expected Smoothness
Kawamoto, Yuta
Iiduka, Hideaki
Machine Learning
Stochastic gradient descent (SGD) is the workhorse of large-scale learning, yet classical analyses rely on assumptions that can be either too strong (bounded variance) or too coarse (uniform noise). The expected smoothness (ES) condition has emerged as a flexible alternative that ties the second moment of stochastic gradients to the objective value and the full gradient. This paper presents a self-contained convergence analysis of SGD under ES. We (i) refine ES with interpretations and sampling-dependent constants; (ii) derive bounds of the expectation of squared full gradient norm; and (iii) prove $O(1/K)$ rates with explicit residual errors for various step-size schedules. All proofs are given in full detail in the appendix. Our treatment unifies and extends recent threads (Khaled and Richtárik, 2020; Umeda and Iiduka, 2025).
title Convergence Analysis of SGD under Expected Smoothness
topic Machine Learning
url https://arxiv.org/abs/2510.20608