Generative Feature Training of Thin 2-Layer Networks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hertrich, Johannes, Neumayer, Sebastian
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916894595874816
author Hertrich, Johannes
Neumayer, Sebastian
author_facet Hertrich, Johannes
Neumayer, Sebastian
contents We consider the approximation of functions by 2-layer neural networks with a small number of hidden weights based on the squared loss and small datasets. Due to the highly non-convex energy landscape, gradient-based training often suffers from local minima. As a remedy, we initialize the hidden weights with samples from a learned proposal distribution, which we parameterize as a deep generative model. To train this model, we exploit the fact that with fixed hidden weights, the optimal output weights solve a linear equation. After learning the generative model, we refine the sampled weights with a gradient-based post-processing in the latent space. Here, we also include a regularization scheme to counteract potential noise. Finally, we demonstrate the effectiveness of our approach by numerical examples.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generative Feature Training of Thin 2-Layer Networks
Hertrich, Johannes
Neumayer, Sebastian
Machine Learning
Numerical Analysis
We consider the approximation of functions by 2-layer neural networks with a small number of hidden weights based on the squared loss and small datasets. Due to the highly non-convex energy landscape, gradient-based training often suffers from local minima. As a remedy, we initialize the hidden weights with samples from a learned proposal distribution, which we parameterize as a deep generative model. To train this model, we exploit the fact that with fixed hidden weights, the optimal output weights solve a linear equation. After learning the generative model, we refine the sampled weights with a gradient-based post-processing in the latent space. Here, we also include a regularization scheme to counteract potential noise. Finally, we demonstrate the effectiveness of our approach by numerical examples.
title Generative Feature Training of Thin 2-Layer Networks
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2411.06848