Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Wenquan, Sui, Yang, Teng, Jiaye, Wang, Bohan, Xu, Jing, Yang, Jingqin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914353142300672
author Ma, Wenquan
Sui, Yang
Teng, Jiaye
Wang, Bohan
Xu, Jing
Yang, Jingqin
author_facet Ma, Wenquan
Sui, Yang
Teng, Jiaye
Wang, Bohan
Xu, Jing
Yang, Jingqin
contents Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize $η_t = \mathcal{O}(1/t)$ under non-convex training regimes, where $t$ denotes iterations. This rigid decay of the stepsize potentially impedes optimization and may not align with practical scenarios. In this paper, we derive the generalization bounds under the homogeneous neural network regimes, proving that this regime enables slower stepsize decay of order $Ω(1/\sqrt{t})$ under mild assumptions. We further extend the theoretical results from several aspects, e.g., non-Lipschitz regimes. This finding is broadly applicable, as homogeneous neural networks encompass fully-connected and convolutional neural networks with ReLU and LeakyReLU activations.
format Preprint
id arxiv_https___arxiv_org_abs_2602_22936
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
Ma, Wenquan
Sui, Yang
Teng, Jiaye
Wang, Bohan
Xu, Jing
Yang, Jingqin
Machine Learning
Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize $η_t = \mathcal{O}(1/t)$ under non-convex training regimes, where $t$ denotes iterations. This rigid decay of the stepsize potentially impedes optimization and may not align with practical scenarios. In this paper, we derive the generalization bounds under the homogeneous neural network regimes, proving that this regime enables slower stepsize decay of order $Ω(1/\sqrt{t})$ under mild assumptions. We further extend the theoretical results from several aspects, e.g., non-Lipschitz regimes. This finding is broadly applicable, as homogeneous neural networks encompass fully-connected and convolutional neural networks with ReLU and LeakyReLU activations.
title Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
topic Machine Learning
url https://arxiv.org/abs/2602.22936