Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Benigni, Lucas, Paquette, Elliot
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914009433767936
author Benigni, Lucas
Paquette, Elliot
author_facet Benigni, Lucas
Paquette, Elliot
contents We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times p}$ is an i.i.d $\mathcal{N}(0,1)$ matrix and $D\in\mathbb{R}^{p\times p}$ is a diagonal matrix with i.i.d bounded entries, we consider the matrix \[ \mathrm{NTK} = \frac{1}{d}XX^\top \odot \frac{1}{p} σ'\left( \frac{1}{\sqrt{d}}XW \right)D^2 σ'\left( \frac{1}{\sqrt{d}}XW \right)^\top \] where $σ'$ is a pseudo-Lipschitz function applied entrywise and under the scaling $\frac{n}{dp}\to γ_1$ and $\frac{p}{d}\to γ_2$. We describe the asymptotic distribution as the free multiplicative convolution of the Marchenko--Pastur distribution with a deterministic distribution depending on $σ$ and $D$.
format Preprint
id arxiv_https___arxiv_org_abs_2508_20036
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
Benigni, Lucas
Paquette, Elliot
Probability
Machine Learning
We compute the asymptotic eigenvalue distribution of the neural tangent kernel of a two-layer neural network under a specific scaling of dimension. Namely, if $X\in\mathbb{R}^{n\times d}$ is an i.i.d random matrix, $W\in\mathbb{R}^{d\times p}$ is an i.i.d $\mathcal{N}(0,1)$ matrix and $D\in\mathbb{R}^{p\times p}$ is a diagonal matrix with i.i.d bounded entries, we consider the matrix \[ \mathrm{NTK} = \frac{1}{d}XX^\top \odot \frac{1}{p} σ'\left( \frac{1}{\sqrt{d}}XW \right)D^2 σ'\left( \frac{1}{\sqrt{d}}XW \right)^\top \] where $σ'$ is a pseudo-Lipschitz function applied entrywise and under the scaling $\frac{n}{dp}\to γ_1$ and $\frac{p}{d}\to γ_2$. We describe the asymptotic distribution as the free multiplicative convolution of the Marchenko--Pastur distribution with a deterministic distribution depending on $σ$ and $D$.
title Eigenvalue distribution of the Neural Tangent Kernel in the quadratic scaling
topic Probability
Machine Learning
url https://arxiv.org/abs/2508.20036