Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lim, Jinwoo, Kim, Suhyun, Moon, Soo-Mook
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917244346302464
author Lim, Jinwoo
Kim, Suhyun
Moon, Soo-Mook
author_facet Lim, Jinwoo
Kim, Suhyun
Moon, Soo-Mook
contents One of the chronic problems of deep-learning models is shortcut learning. In a case where the majority of training data are dominated by a certain feature, neural networks prefer to learn such a feature even if the feature is not generalizable outside the training set. Based on the framework of Neural Tangent Kernel (NTK), we analyzed the case of linear neural networks to derive some important properties of shortcut learning. We defined a feature of a neural network as an eigenfunction of NTK. Then, we found that shortcut features correspond to features with larger eigenvalues when the shortcuts stem from the imbalanced number of samples in the clustered distribution. We also showed that the features with larger eigenvalues still have a large influence on the neural network output even after training, due to data variances in the clusters. Such a preference for certain features remains even when a margin of a neural network output is controlled, which shows that the max-margin bias is not the only major reason for shortcut learning. These properties of linear neural networks are empirically extended for more complex neural networks as a two-layer fully-connected ReLU network and a ResNet-18.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03066
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More
Lim, Jinwoo
Kim, Suhyun
Moon, Soo-Mook
Machine Learning
Artificial Intelligence
One of the chronic problems of deep-learning models is shortcut learning. In a case where the majority of training data are dominated by a certain feature, neural networks prefer to learn such a feature even if the feature is not generalizable outside the training set. Based on the framework of Neural Tangent Kernel (NTK), we analyzed the case of linear neural networks to derive some important properties of shortcut learning. We defined a feature of a neural network as an eigenfunction of NTK. Then, we found that shortcut features correspond to features with larger eigenvalues when the shortcuts stem from the imbalanced number of samples in the clustered distribution. We also showed that the features with larger eigenvalues still have a large influence on the neural network output even after training, due to data variances in the clusters. Such a preference for certain features remains even when a margin of a neural network output is controlled, which shows that the max-margin bias is not the only major reason for shortcut learning. These properties of linear neural networks are empirically extended for more complex neural networks as a two-layer fully-connected ReLU network and a ResNet-18.
title Shortcut Features as Top Eigenfunctions of NTK: A Linear Neural Network Case and More
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03066