Training NTK to Generalize with KARE

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schwab, Johannes, Kelly, Bryan, Malamud, Semyon, Xu, Teng Andrea
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909618604605440
author Schwab, Johannes
Kelly, Bryan
Malamud, Semyon
Xu, Teng Andrea
author_facet Schwab, Johannes
Kelly, Bryan
Malamud, Semyon
Xu, Teng Andrea
contents The performance of the data-dependent neural tangent kernel (NTK; Jacot et al. (2018)) associated with a trained deep neural network (DNN) often matches or exceeds that of the full network. This implies that DNN training via gradient descent implicitly performs kernel learning by optimizing the NTK. In this paper, we propose instead to optimize the NTK explicitly. Rather than minimizing empirical risk, we train the NTK to minimize its generalization error using the recently developed Kernel Alignment Risk Estimator (KARE; Jacot et al. (2020)). Our simulations and real data experiments show that NTKs trained with KARE consistently match or significantly outperform the original DNN and the DNN- induced NTK (the after-kernel). These results suggest that explicitly trained kernels can outperform traditional end-to-end DNN optimization in certain settings, challenging the conventional dominance of DNNs. We argue that explicit training of NTK is a form of over-parametrized feature learning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11347
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training NTK to Generalize with KARE
Schwab, Johannes
Kelly, Bryan
Malamud, Semyon
Xu, Teng Andrea
Machine Learning
The performance of the data-dependent neural tangent kernel (NTK; Jacot et al. (2018)) associated with a trained deep neural network (DNN) often matches or exceeds that of the full network. This implies that DNN training via gradient descent implicitly performs kernel learning by optimizing the NTK. In this paper, we propose instead to optimize the NTK explicitly. Rather than minimizing empirical risk, we train the NTK to minimize its generalization error using the recently developed Kernel Alignment Risk Estimator (KARE; Jacot et al. (2020)). Our simulations and real data experiments show that NTKs trained with KARE consistently match or significantly outperform the original DNN and the DNN- induced NTK (the after-kernel). These results suggest that explicitly trained kernels can outperform traditional end-to-end DNN optimization in certain settings, challenging the conventional dominance of DNNs. We argue that explicit training of NTK is a form of over-parametrized feature learning.
title Training NTK to Generalize with KARE
topic Machine Learning
url https://arxiv.org/abs/2505.11347