Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahverdi, Vahid, Marchetti, Giovanni Luca, Kohn, Kathlén
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910020694704128
author Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
author_facet Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
contents We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Networks (CNNs) with an activation function that is a sufficiently generic polynomial. First, we address the identifiability problem, showing that, for almost all functions in the neuromanifold of an MLP, there exist only finitely many parameter choices yielding that function. For CNNs, the parametrization is generically one-to-one. As a consequence, we compute the dimension of the neuromanifold. Second, we describe singular points of neuromanifolds. We characterize singularities completely for CNNs, and partially for MLPs. In both cases, they arise from sparse subnetworks. For MLPs, we prove that these singularities often correspond to critical points of the mean-squared error loss, which does not hold for CNNs. This provides a geometric explanation of the sparsity bias of MLPs. All of our results leverage tools from algebraic geometry.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11846
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
Shahverdi, Vahid
Marchetti, Giovanni Luca
Kohn, Kathlén
Machine Learning
Algebraic Geometry
We study function spaces parametrized by neural networks, referred to as neuromanifolds. Specifically, we focus on deep Multi-Layer Perceptrons (MLPs) and Convolutional Neural Networks (CNNs) with an activation function that is a sufficiently generic polynomial. First, we address the identifiability problem, showing that, for almost all functions in the neuromanifold of an MLP, there exist only finitely many parameter choices yielding that function. For CNNs, the parametrization is generically one-to-one. As a consequence, we compute the dimension of the neuromanifold. Second, we describe singular points of neuromanifolds. We characterize singularities completely for CNNs, and partially for MLPs. In both cases, they arise from sparse subnetworks. For MLPs, we prove that these singularities often correspond to critical points of the mean-squared error loss, which does not hold for CNNs. This provides a geometric explanation of the sparsity bias of MLPs. All of our results leverage tools from algebraic geometry.
title Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural Networks
topic Machine Learning
Algebraic Geometry
url https://arxiv.org/abs/2505.11846