From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rubin, Noa, Fischer, Kirsten, Lindner, Javed, Dahmen, David, Seroussi, Inbar, Ringel, Zohar, Krämer, Michael, Helias, Moritz
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918036990066688
author Rubin, Noa
Fischer, Kirsten
Lindner, Javed
Dahmen, David
Seroussi, Inbar
Ringel, Zohar
Krämer, Michael
Helias, Moritz
author_facet Rubin, Noa
Fischer, Kirsten
Lindner, Javed
Dahmen, David
Seroussi, Inbar
Ringel, Zohar
Krämer, Michael
Helias, Moritz
contents Feature learning in neural networks is crucial for their expressive power and inductive biases, motivating various theoretical approaches. Some approaches describe network behavior after training through a change in kernel scale from initialization, resulting in a generalization power comparable to a Gaussian process. Conversely, in other approaches training results in the adaptation of the kernel to the data, involving directional changes to the kernel. The relationship and respective strengths of these two views have so far remained unresolved. This work presents a theoretical framework of multi-scale adaptive feature learning bridging these two views. Using methods from statistical mechanics, we derive analytical expressions for network output statistics which are valid across scaling regimes and in the continuum between them. A systematic expansion of the network's probability distribution reveals that mean-field scaling requires only a saddle-point approximation, while standard scaling necessitates additional correction terms. Remarkably, we find across regimes that kernel adaptation can be reduced to an effective kernel rescaling when predicting the mean network output in the special case of a linear network. However, for linear and non-linear networks, the multi-scale adaptive approach captures directional feature learning effects, providing richer insights than what could be recovered from a rescaling of the kernel alone.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03210
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
Rubin, Noa
Fischer, Kirsten
Lindner, Javed
Dahmen, David
Seroussi, Inbar
Ringel, Zohar
Krämer, Michael
Helias, Moritz
Disordered Systems and Neural Networks
Machine Learning
Feature learning in neural networks is crucial for their expressive power and inductive biases, motivating various theoretical approaches. Some approaches describe network behavior after training through a change in kernel scale from initialization, resulting in a generalization power comparable to a Gaussian process. Conversely, in other approaches training results in the adaptation of the kernel to the data, involving directional changes to the kernel. The relationship and respective strengths of these two views have so far remained unresolved. This work presents a theoretical framework of multi-scale adaptive feature learning bridging these two views. Using methods from statistical mechanics, we derive analytical expressions for network output statistics which are valid across scaling regimes and in the continuum between them. A systematic expansion of the network's probability distribution reveals that mean-field scaling requires only a saddle-point approximation, while standard scaling necessitates additional correction terms. Remarkably, we find across regimes that kernel adaptation can be reduced to an effective kernel rescaling when predicting the mean network output in the special case of a linear network. However, for linear and non-linear networks, the multi-scale adaptive approach captures directional feature learning effects, providing richer insights than what could be recovered from a rescaling of the kernel alone.
title From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning
topic Disordered Systems and Neural Networks
Machine Learning
url https://arxiv.org/abs/2502.03210