Gradient Flow Convergence Guarantee for General Neural Network Architectures

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Jakhmola, Yash
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912612921376768
author Jakhmola, Yash
author_facet Jakhmola, Yash
contents A key challenge in modern deep learning theory is to explain the remarkable success of gradient-based optimization methods when training large-scale, complex deep neural networks. Though linear convergence of such methods has been proved for a handful of specific architectures, a united theory still evades researchers. This article presents a unified proof for linear convergence of continuous gradient descent, also called gradient flow, while training any neural network with piecewise non-zero polynomial activations or ReLU, sigmoid activations. Our primary contribution is a single, general theorem that not only covers architectures for which this result was previously unknown but also consolidates existing results under weaker assumptions. While our focus is theoretical and our results are only exact in the infinitesimal step size limit, we nevertheless find excellent empirical agreement between the predictions of our result and those of the practical step-size gradient descent method.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient Flow Convergence Guarantee for General Neural Network Architectures
Jakhmola, Yash
Machine Learning
Artificial Intelligence
A key challenge in modern deep learning theory is to explain the remarkable success of gradient-based optimization methods when training large-scale, complex deep neural networks. Though linear convergence of such methods has been proved for a handful of specific architectures, a united theory still evades researchers. This article presents a unified proof for linear convergence of continuous gradient descent, also called gradient flow, while training any neural network with piecewise non-zero polynomial activations or ReLU, sigmoid activations. Our primary contribution is a single, general theorem that not only covers architectures for which this result was previously unknown but also consolidates existing results under weaker assumptions. While our focus is theoretical and our results are only exact in the infinitesimal step size limit, we nevertheless find excellent empirical agreement between the predictions of our result and those of the practical step-size gradient descent method.
title Gradient Flow Convergence Guarantee for General Neural Network Architectures
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2509.23887