Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866914353142300672 |
|---|---|
| author | Ma, Wenquan Sui, Yang Teng, Jiaye Wang, Bohan Xu, Jing Yang, Jingqin |
| author_facet | Ma, Wenquan Sui, Yang Teng, Jiaye Wang, Bohan Xu, Jing Yang, Jingqin |
| contents | Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize $η_t = \mathcal{O}(1/t)$ under non-convex training regimes, where $t$ denotes iterations. This rigid decay of the stepsize potentially impedes optimization and may not align with practical scenarios. In this paper, we derive the generalization bounds under the homogeneous neural network regimes, proving that this regime enables slower stepsize decay of order $Ω(1/\sqrt{t})$ under mild assumptions. We further extend the theoretical results from several aspects, e.g., non-Lipschitz regimes. This finding is broadly applicable, as homogeneous neural networks encompass fully-connected and convolutional neural networks with ReLU and LeakyReLU activations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_22936 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks Ma, Wenquan Sui, Yang Teng, Jiaye Wang, Bohan Xu, Jing Yang, Jingqin Machine Learning Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize $η_t = \mathcal{O}(1/t)$ under non-convex training regimes, where $t$ denotes iterations. This rigid decay of the stepsize potentially impedes optimization and may not align with practical scenarios. In this paper, we derive the generalization bounds under the homogeneous neural network regimes, proving that this regime enables slower stepsize decay of order $Ω(1/\sqrt{t})$ under mild assumptions. We further extend the theoretical results from several aspects, e.g., non-Lipschitz regimes. This finding is broadly applicable, as homogeneous neural networks encompass fully-connected and convolutional neural networks with ReLU and LeakyReLU activations. |
| title | Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2602.22936 |