A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality
Fuente:
arXiv
Guardado en:
| Autor principal: | |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911838176804864 |
|---|---|
| author | Isobe, Noboru |
| author_facet | Isobe, Noboru |
| contents | This study focuses on a Wasserstein-type gradient flow, which represents an optimization process of a continuous model of a Deep Neural Network (DNN). First, we establish the existence of a minimizer for an average loss of the model under $L^2$-regularization. Subsequently, we show the existence of a curve of maximal slope of the loss. Our main result is the convergence of flow to a critical point of the loss as time goes to infinity. An essential aspect of proving this result involves the establishment of the Łojasiewicz--Simon gradient inequality for the loss. We derive this inequality by assuming the analyticity of NNs and loss functions. Our proofs offer a new approach for analyzing the asymptotic behavior of Wasserstein-type gradient flows for nonconvex functionals. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_15365 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality Isobe, Noboru Machine Learning Analysis of PDEs Functional Analysis Probability 35B40, 49J20, 49Q22, 68T07 This study focuses on a Wasserstein-type gradient flow, which represents an optimization process of a continuous model of a Deep Neural Network (DNN). First, we establish the existence of a minimizer for an average loss of the model under $L^2$-regularization. Subsequently, we show the existence of a curve of maximal slope of the loss. Our main result is the convergence of flow to a critical point of the loss as time goes to infinity. An essential aspect of proving this result involves the establishment of the Łojasiewicz--Simon gradient inequality for the loss. We derive this inequality by assuming the analyticity of NNs and loss functions. Our proofs offer a new approach for analyzing the asymptotic behavior of Wasserstein-type gradient flows for nonconvex functionals. |
| title | A convergence result of a continuous model of deep learning via Łojasiewicz--Simon inequality |
| topic | Machine Learning Analysis of PDEs Functional Analysis Probability 35B40, 49J20, 49Q22, 68T07 |
| url | https://arxiv.org/abs/2311.15365 |