Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics
Fuente:
arXiv
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909562773176320 |
|---|---|
| author | Liu, Kang Zuazua, Enrique |
| author_facet | Liu, Kang Zuazua, Enrique |
| contents | In this work, we address three non-convex optimization problems associated with the training of shallow neural networks (NNs) for exact and approximate representation, as well as for regression tasks. Through a mean-field approach, we convexify these problems and, applying a representer theorem, prove the absence of relaxation gaps. We establish generalization bounds for the resulting NN solutions, assessing their predictive performance on test datasets and, analyzing the impact of key hyperparameters on these bounds, propose optimal choices.
On the computational side, we examine the discretization of the convexified problems and derive convergence rates. For low-dimensional datasets, these discretized problems are efficiently solvable using the simplex method. For high-dimensional datasets, we propose a sparsification algorithm that, combined with gradient descent for over-parameterized shallow NNs, yields effective solutions to the primal problems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_01619 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics Liu, Kang Zuazua, Enrique Machine Learning Optimization and Control 68T07, 68T09, 90C06, 90C26 In this work, we address three non-convex optimization problems associated with the training of shallow neural networks (NNs) for exact and approximate representation, as well as for regression tasks. Through a mean-field approach, we convexify these problems and, applying a representer theorem, prove the absence of relaxation gaps. We establish generalization bounds for the resulting NN solutions, assessing their predictive performance on test datasets and, analyzing the impact of key hyperparameters on these bounds, propose optimal choices. On the computational side, we examine the discretization of the convexified problems and derive convergence rates. For low-dimensional datasets, these discretized problems are efficiently solvable using the simplex method. For high-dimensional datasets, we propose a sparsification algorithm that, combined with gradient descent for over-parameterized shallow NNs, yields effective solutions to the primal problems. |
| title | Representation and Regression Problems in Neural Networks: Relaxation, Generalization, and Numerics |
| topic | Machine Learning Optimization and Control 68T07, 68T09, 90C06, 90C26 |
| url | https://arxiv.org/abs/2412.01619 |