Approximation and Gradient Descent Training with Neural Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Welper, G.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914802628034560
author Welper, G.
author_facet Welper, G.
contents It is well understood that neural networks with carefully hand-picked weights provide powerful function approximation and that they can be successfully trained in over-parametrized regimes. Since over-parametrization ensures zero training error, these two theories are not immediately compatible. Recent work uses the smoothness that is required for approximation results to extend a neural tangent kernel (NTK) optimization argument to an under-parametrized regime and show direct approximation bounds for networks trained by gradient flow. Since gradient flow is only an idealization of a practical method, this paper establishes analogous results for networks trained by gradient descent.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11696
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Approximation and Gradient Descent Training with Neural Networks
Welper, G.
Machine Learning
41A46, 65K10, 68T07
It is well understood that neural networks with carefully hand-picked weights provide powerful function approximation and that they can be successfully trained in over-parametrized regimes. Since over-parametrization ensures zero training error, these two theories are not immediately compatible. Recent work uses the smoothness that is required for approximation results to extend a neural tangent kernel (NTK) optimization argument to an under-parametrized regime and show direct approximation bounds for networks trained by gradient flow. Since gradient flow is only an idealization of a practical method, this paper establishes analogous results for networks trained by gradient descent.
title Approximation and Gradient Descent Training with Neural Networks
topic Machine Learning
41A46, 65K10, 68T07
url https://arxiv.org/abs/2405.11696