The Infinite Model
Fuente:
Zenodo
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Recurso digital |
| Lingua: | inglese |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902153754312704 |
|---|---|
| author | Hobe-Gelting, Victor da Rocha, Paulo M. M. |
| author_facet | Hobe-Gelting, Victor da Rocha, Paulo M. M. |
| contents | <p>Overparameterized neural networks exhibit a persistent simplicity bias: among the many functions that interpolate the data, the induced function law preferentially selects simpler solutions. We give a unified theory of this bias by identifying the geometric quantity that governs selection at finite width and its explicit limit at infinite width. At finite width, the induced prior on functions is controlled by the minimum parameter cost required to realize a function, yielding an exact selection law and a refined description of relative preference among interpolants. At infinite width, whenever the pushforward prior converges to a Gaussian process, the governing potential admits a closed form determined by the reproducing-kernel geometry of the limiting model. A bridge theorem connects the finite-width and infinite-width theories, showing that infinite-width simplicity bias arises as the limit of finite-width geometry rather than as a separate principle. In the lazy regime, this yields a single equilibrium framework for minimum-norm interpolation, PAC-Bayes complexity, barrier suppression, and scaling behavior. Beyond the fixed-kernel setting, we formulate the kernel-adaptive potential for feature learning, prove its cylindrical GPS description unconditionally, establish the full function-space lift for compact kernel families and for single-hidden-layer networks, and show that the non-lazy trajectory is not closed in kernel - function variables; the remaining architecture-specific task is the coercivity verification for general deep architectures.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20395890 |
| institution | Zenodo |
| language | eng |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | The Infinite Model Hobe-Gelting, Victor da Rocha, Paulo M. M. simplicity bias overparameterized neural networks large deviation principle RKHS Gaussian process NTK infinite-width limit PAC-Bayes scaling laws architecture fading feature learning Onsager-Machlup functional deep learning neural network theory machine learning overparameterization <p>Overparameterized neural networks exhibit a persistent simplicity bias: among the many functions that interpolate the data, the induced function law preferentially selects simpler solutions. We give a unified theory of this bias by identifying the geometric quantity that governs selection at finite width and its explicit limit at infinite width. At finite width, the induced prior on functions is controlled by the minimum parameter cost required to realize a function, yielding an exact selection law and a refined description of relative preference among interpolants. At infinite width, whenever the pushforward prior converges to a Gaussian process, the governing potential admits a closed form determined by the reproducing-kernel geometry of the limiting model. A bridge theorem connects the finite-width and infinite-width theories, showing that infinite-width simplicity bias arises as the limit of finite-width geometry rather than as a separate principle. In the lazy regime, this yields a single equilibrium framework for minimum-norm interpolation, PAC-Bayes complexity, barrier suppression, and scaling behavior. Beyond the fixed-kernel setting, we formulate the kernel-adaptive potential for feature learning, prove its cylindrical GPS description unconditionally, establish the full function-space lift for compact kernel families and for single-hidden-layer networks, and show that the non-lazy trajectory is not closed in kernel - function variables; the remaining architecture-specific task is the coercivity verification for general deep architectures.</p> |
| title | The Infinite Model |
| topic | simplicity bias overparameterized neural networks large deviation principle RKHS Gaussian process NTK infinite-width limit PAC-Bayes scaling laws architecture fading feature learning Onsager-Machlup functional deep learning neural network theory machine learning overparameterization |
| url | https://doi.org/10.5281/zenodo.20395890 |