On the Geometry and Optimization of Polynomial Convolutional Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shahverdi, Vahid, Marchetti, Giovanni Luca, Kohn, Kathlén
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912256428605440
author Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
author_facet Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
contents We study convolutional neural networks with monomial activation functions. Specifically, we prove that their parameterization map is regular and is an isomorphism almost everywhere, up to rescaling the filters. By leveraging on tools from algebraic geometry, we explore the geometric properties of the image in function space of this map - typically referred to as neuromanifold. In particular, we compute the dimension and the degree of the neuromanifold, which measure the expressivity of the model, and describe its singularities. Moreover, for a generic large dataset, we derive an explicit formula that quantifies the number of critical points arising in the optimization of a regression loss.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Geometry and Optimization of Polynomial Convolutional Networks
Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
Machine Learning
Algebraic Geometry
We study convolutional neural networks with monomial activation functions. Specifically, we prove that their parameterization map is regular and is an isomorphism almost everywhere, up to rescaling the filters. By leveraging on tools from algebraic geometry, we explore the geometric properties of the image in function space of this map - typically referred to as neuromanifold. In particular, we compute the dimension and the degree of the neuromanifold, which measure the expressivity of the model, and describe its singularities. Moreover, for a generic large dataset, we derive an explicit formula that quantifies the number of critical points arising in the optimization of a regression loss.
title On the Geometry and Optimization of Polynomial Convolutional Networks
topic Machine Learning
Algebraic Geometry
url https://arxiv.org/abs/2410.00722