Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yilan, Wang, Zhichao, Huang, Wei, Han, Andi, Suzuki, Taiji, Mazumdar, Arya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916792625004544
author Chen, Yilan
Wang, Zhichao
Huang, Wei
Han, Andi
Suzuki, Taiji
Mazumdar, Arya
author_facet Chen, Yilan
Wang, Zhichao
Huang, Wei
Han, Andi
Suzuki, Taiji
Mazumdar, Arya
contents Gradient-based optimization methods have shown remarkable empirical success, yet their theoretical generalization properties remain only partially understood. In this paper, we establish a generalization bound for gradient flow that aligns with the classical Rademacher complexity bounds for kernel methods-specifically those based on the RKHS norm and kernel trace-through a data-dependent kernel called the loss path kernel (LPK). Unlike static kernels such as NTK, the LPK captures the entire training trajectory, adapting to both data and optimization dynamics, leading to tighter and more informative generalization guarantees. Moreover, the bound highlights how the norm of the training loss gradients along the optimization trajectory influences the final generalization performance. The key technical ingredients in our proof combine stability analysis of gradient flow with uniform convergence via Rademacher complexity. Our bound recovers existing kernel regression bounds for overparameterized neural networks and shows the feature learning capability of neural networks compared to kernel methods. Numerical experiments on real-world datasets validate that our bounds correlate well with the true generalization gap.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11357
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
Chen, Yilan
Wang, Zhichao
Huang, Wei
Han, Andi
Suzuki, Taiji
Mazumdar, Arya
Machine Learning
Gradient-based optimization methods have shown remarkable empirical success, yet their theoretical generalization properties remain only partially understood. In this paper, we establish a generalization bound for gradient flow that aligns with the classical Rademacher complexity bounds for kernel methods-specifically those based on the RKHS norm and kernel trace-through a data-dependent kernel called the loss path kernel (LPK). Unlike static kernels such as NTK, the LPK captures the entire training trajectory, adapting to both data and optimization dynamics, leading to tighter and more informative generalization guarantees. Moreover, the bound highlights how the norm of the training loss gradients along the optimization trajectory influences the final generalization performance. The key technical ingredients in our proof combine stability analysis of gradient flow with uniform convergence via Rademacher complexity. Our bound recovers existing kernel regression bounds for overparameterized neural networks and shows the feature learning capability of neural networks compared to kernel methods. Numerical experiments on real-world datasets validate that our bounds correlate well with the true generalization gap.
title Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
topic Machine Learning
url https://arxiv.org/abs/2506.11357