Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hölzl, Florian A., Rueckert, Daniel, Kaissis, Georgios
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911239508066304
author Hölzl, Florian A.
Rueckert, Daniel
Kaissis, Georgios
author_facet Hölzl, Florian A.
Rueckert, Daniel
Kaissis, Georgios
contents Robust validation metrics remain essential in contemporary deep learning, not only to detect overfitting and poor generalization, but also to monitor training dynamics. In the supervised classification setting, we investigate whether interactions between training data and model weights can yield such a metric that both tracks generalization during training and attributes performance to individual training samples. We introduce Gradient-Weight Alignment (GWA), quantifying the coherence between per-sample gradients and model weights. We show that effective learning corresponds to coherent alignment, while misalignment indicates deteriorating generalization. GWA is efficiently computable during training and reflects both sample-specific contributions and dataset-wide learning dynamics. Extensive experiments show that GWA accurately predicts optimal early stopping, enables principled model comparisons, and identifies influential training samples, providing a validation-set-free approach for model analysis directly from the training data.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25480
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
Hölzl, Florian A.
Rueckert, Daniel
Kaissis, Georgios
Machine Learning
Robust validation metrics remain essential in contemporary deep learning, not only to detect overfitting and poor generalization, but also to monitor training dynamics. In the supervised classification setting, we investigate whether interactions between training data and model weights can yield such a metric that both tracks generalization during training and attributes performance to individual training samples. We introduce Gradient-Weight Alignment (GWA), quantifying the coherence between per-sample gradients and model weights. We show that effective learning corresponds to coherent alignment, while misalignment indicates deteriorating generalization. GWA is efficiently computable during training and reflects both sample-specific contributions and dataset-wide learning dynamics. Extensive experiments show that GWA accurately predicts optimal early stopping, enables principled model comparisons, and identifies influential training samples, providing a validation-set-free approach for model analysis directly from the training data.
title Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
topic Machine Learning
url https://arxiv.org/abs/2510.25480