The Infinite Model

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Hobe-Gelting, Victor, da Rocha, Paulo M. M.
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902153754312704
author Hobe-Gelting, Victor
da Rocha, Paulo M. M.
author_facet Hobe-Gelting, Victor
da Rocha, Paulo M. M.
contents <p>Overparameterized neural networks exhibit a persistent simplicity bias: among the many functions that interpolate the data, the induced function law preferentially selects simpler solutions. We give a unified theory of this bias by identifying the geometric quantity that governs selection at finite width and its explicit limit at infinite width. At finite width, the induced prior on functions is controlled by the minimum parameter cost required to realize a function, yielding an exact selection law and a refined description of relative preference among interpolants. At infinite width, whenever the pushforward prior converges to a Gaussian process, the governing potential admits a closed form determined by the reproducing-kernel geometry of the limiting model. A bridge theorem connects the finite-width and infinite-width theories, showing that infinite-width simplicity bias arises as the limit of finite-width geometry rather than as a separate principle. In the lazy regime, this yields a single equilibrium framework for minimum-norm interpolation, PAC-Bayes complexity, barrier suppression, and scaling behavior. Beyond the fixed-kernel setting, we formulate the kernel-adaptive potential for feature learning, prove its cylindrical GPS description unconditionally, establish the full function-space lift for compact kernel families and for single-hidden-layer networks, and show that the non-lazy trajectory is not closed in kernel - function variables; the remaining architecture-specific task is the coercivity verification for general deep architectures.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20395890
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle The Infinite Model
Hobe-Gelting, Victor
da Rocha, Paulo M. M.
simplicity bias
overparameterized neural networks
large deviation principle
RKHS
Gaussian process
NTK
infinite-width limit
PAC-Bayes
scaling laws
architecture fading
feature learning
Onsager-Machlup functional
deep learning
neural network theory
machine learning
overparameterization
<p>Overparameterized neural networks exhibit a persistent simplicity bias: among the many functions that interpolate the data, the induced function law preferentially selects simpler solutions. We give a unified theory of this bias by identifying the geometric quantity that governs selection at finite width and its explicit limit at infinite width. At finite width, the induced prior on functions is controlled by the minimum parameter cost required to realize a function, yielding an exact selection law and a refined description of relative preference among interpolants. At infinite width, whenever the pushforward prior converges to a Gaussian process, the governing potential admits a closed form determined by the reproducing-kernel geometry of the limiting model. A bridge theorem connects the finite-width and infinite-width theories, showing that infinite-width simplicity bias arises as the limit of finite-width geometry rather than as a separate principle. In the lazy regime, this yields a single equilibrium framework for minimum-norm interpolation, PAC-Bayes complexity, barrier suppression, and scaling behavior. Beyond the fixed-kernel setting, we formulate the kernel-adaptive potential for feature learning, prove its cylindrical GPS description unconditionally, establish the full function-space lift for compact kernel families and for single-hidden-layer networks, and show that the non-lazy trajectory is not closed in kernel - function variables; the remaining architecture-specific task is the coercivity verification for general deep architectures.</p>
title The Infinite Model
topic simplicity bias
overparameterized neural networks
large deviation principle
RKHS
Gaussian process
NTK
infinite-width limit
PAC-Bayes
scaling laws
architecture fading
feature learning
Onsager-Machlup functional
deep learning
neural network theory
machine learning
overparameterization
url https://doi.org/10.5281/zenodo.20395890