A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Isobe, Noboru
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911838176804864
author Isobe, Noboru
author_facet Isobe, Noboru
contents This study focuses on a Wasserstein-type gradient flow, which represents an optimization process of a continuous model of a Deep Neural Network (DNN). First, we establish the existence of a minimizer for an average loss of the model under $L^2$-regularization. Subsequently, we show the existence of a curve of maximal slope of the loss. Our main result is the convergence of flow to a critical point of the loss as time goes to infinity. An essential aspect of proving this result involves the establishment of the Łojasiewicz--Simon gradient inequality for the loss. We derive this inequality by assuming the analyticity of NNs and loss functions. Our proofs offer a new approach for analyzing the asymptotic behavior of Wasserstein-type gradient flows for nonconvex functionals.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15365
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality
Isobe, Noboru
Machine Learning
Analysis of PDEs
Functional Analysis
Probability
35B40, 49J20, 49Q22, 68T07
This study focuses on a Wasserstein-type gradient flow, which represents an optimization process of a continuous model of a Deep Neural Network (DNN). First, we establish the existence of a minimizer for an average loss of the model under $L^2$-regularization. Subsequently, we show the existence of a curve of maximal slope of the loss. Our main result is the convergence of flow to a critical point of the loss as time goes to infinity. An essential aspect of proving this result involves the establishment of the Łojasiewicz--Simon gradient inequality for the loss. We derive this inequality by assuming the analyticity of NNs and loss functions. Our proofs offer a new approach for analyzing the asymptotic behavior of Wasserstein-type gradient flows for nonconvex functionals.
title A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality
topic Machine Learning
Analysis of PDEs
Functional Analysis
Probability
35B40, 49J20, 49Q22, 68T07
url https://arxiv.org/abs/2311.15365