Understanding the role of depth in the neural tangent kernel for overparameterized neural networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: St-Arnaud, William, Carvalho, Margarida, Farnadi, Golnoosh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911258021724160
author St-Arnaud, William
Carvalho, Margarida
Farnadi, Golnoosh
author_facet St-Arnaud, William
Carvalho, Margarida
Farnadi, Golnoosh
contents Overparameterized fully-connected neural networks have been shown to behave like kernel models when trained with gradient descent, under mild conditions on the width, the learning rate, and the parameter initialization. In the limit of infinitely large widths and small learning rate, the kernel that is obtained allows to represent the output of the learned model with a closed-form solution. This closed-form solution hinges on the invertibility of the limiting kernel, a property that often holds on real-world datasets. In this work, we analyze the sensitivity of large ReLU networks to increasing depths by characterizing the corresponding limiting kernel. Our theoretical results demonstrate that the normalized limiting kernel approaches the matrix of ones. In contrast, they show the corresponding closed-form solution approaches a fixed limit on the sphere. We empirically evaluate the order of magnitude in network depth required to observe this convergent behavior, and we describe the essential properties that enable the generalization of our results to other kernels.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07272
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the role of depth in the neural tangent kernel for overparameterized neural networks
St-Arnaud, William
Carvalho, Margarida
Farnadi, Golnoosh
Machine Learning
Overparameterized fully-connected neural networks have been shown to behave like kernel models when trained with gradient descent, under mild conditions on the width, the learning rate, and the parameter initialization. In the limit of infinitely large widths and small learning rate, the kernel that is obtained allows to represent the output of the learned model with a closed-form solution. This closed-form solution hinges on the invertibility of the limiting kernel, a property that often holds on real-world datasets. In this work, we analyze the sensitivity of large ReLU networks to increasing depths by characterizing the corresponding limiting kernel. Our theoretical results demonstrate that the normalized limiting kernel approaches the matrix of ones. In contrast, they show the corresponding closed-form solution approaches a fixed limit on the sphere. We empirically evaluate the order of magnitude in network depth required to observe this convergent behavior, and we describe the essential properties that enable the generalization of our results to other kernels.
title Understanding the role of depth in the neural tangent kernel for overparameterized neural networks
topic Machine Learning
url https://arxiv.org/abs/2511.07272