How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wei, Han, Andi, Song, Yujin, Chen, Yilan, Wu, Denny, Zou, Difan, Suzuki, Taiji
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917027971596288
author Huang, Wei
Han, Andi
Song, Yujin
Chen, Yilan
Wu, Denny
Zou, Difan
Suzuki, Taiji
author_facet Huang, Wei
Han, Andi
Song, Yujin
Chen, Yilan
Wu, Denny
Zou, Difan
Suzuki, Taiji
contents The capacity of deep learning models is often large enough to both learn the underlying statistical signal and overfit to noise in the training set. This noise memorization can be harmful especially for data with a low signal-to-noise ratio (SNR), leading to poor generalization. Inspired by prior observations that label noise provides implicit regularization that improves generalization, in this work, we investigate whether introducing label noise to the gradient updates can enhance the test performance of neural network (NN) in the low SNR regime. Specifically, we consider training a two-layer NN with a simple label noise gradient descent (GD) algorithm, in an idealized signal-noise data setting. We prove that adding label noise during training suppresses noise memorization, preventing it from dominating the learning process; consequently, label noise GD enjoys rapid signal growth while the overfitting remains controlled, thereby achieving good generalization despite the low SNR. In contrast, we also show that NN trained with standard GD tends to overfit to noise in the same low SNR setting and establish a non-vanishing lower bound on its test error, thus demonstrating the benefit of introducing label noise in gradient-based training.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17526
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
Huang, Wei
Han, Andi
Song, Yujin
Chen, Yilan
Wu, Denny
Zou, Difan
Suzuki, Taiji
Machine Learning
The capacity of deep learning models is often large enough to both learn the underlying statistical signal and overfit to noise in the training set. This noise memorization can be harmful especially for data with a low signal-to-noise ratio (SNR), leading to poor generalization. Inspired by prior observations that label noise provides implicit regularization that improves generalization, in this work, we investigate whether introducing label noise to the gradient updates can enhance the test performance of neural network (NN) in the low SNR regime. Specifically, we consider training a two-layer NN with a simple label noise gradient descent (GD) algorithm, in an idealized signal-noise data setting. We prove that adding label noise during training suppresses noise memorization, preventing it from dominating the learning process; consequently, label noise GD enjoys rapid signal growth while the overfitting remains controlled, thereby achieving good generalization despite the low SNR. In contrast, we also show that NN trained with standard GD tends to overfit to noise in the same low SNR setting and establish a non-vanishing lower bound on its test error, thus demonstrating the benefit of introducing label noise in gradient-based training.
title How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
topic Machine Learning
url https://arxiv.org/abs/2510.17526