Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Caron, Francois, Ayed, Fadhel, Jung, Paul, Lee, Hoil, Lee, Juho, Yang, Hongseok
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909498065551360
author Caron, Francois
Ayed, Fadhel
Jung, Paul
Lee, Hoil
Lee, Juho
Yang, Hongseok
author_facet Caron, Francois
Ayed, Fadhel
Jung, Paul
Lee, Hoil
Lee, Juho
Yang, Hongseok
contents We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel (NTK) parameterisation. We prove that for large such neural networks, with high probability, gradient flow and gradient descent converge to a global minimum and can learn features in some sense, unlike in the NTK parameterisation. We perform experiments illustrating our theoretical results and discuss the benefits of such scaling in terms of prunability and transfer learning.
format Preprint
id arxiv_https___arxiv_org_abs_2302_01002
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning
Caron, Francois
Ayed, Fadhel
Jung, Paul
Lee, Hoil
Lee, Juho
Yang, Hongseok
Machine Learning
Optimization and Control
We consider gradient-based optimisation of wide, shallow neural networks, where the output of each hidden node is scaled by a positive parameter. The scaling parameters are non-identical, differing from the classical Neural Tangent Kernel (NTK) parameterisation. We prove that for large such neural networks, with high probability, gradient flow and gradient descent converge to a global minimum and can learn features in some sense, unlike in the NTK parameterisation. We perform experiments illustrating our theoretical results and discuss the benefits of such scaling in terms of prunability and transfer learning.
title Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning
topic Machine Learning
Optimization and Control
url https://arxiv.org/abs/2302.01002