Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sanyal, Amartya, Hu, Yaxi, Yu, Yaodong, Ma, Yian, Wang, Yixin, Schölkopf, Bernhard
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929401881427968
author Sanyal, Amartya
Hu, Yaxi
Yu, Yaodong
Ma, Yian
Wang, Yixin
Schölkopf, Bernhard
author_facet Sanyal, Amartya
Hu, Yaxi
Yu, Yaodong
Ma, Yian
Wang, Yixin
Schölkopf, Bernhard
contents "Accuracy-on-the-line" is a widely observed phenomenon in machine learning, where a model's accuracy on in-distribution (ID) and out-of-distribution (OOD) data is positively correlated across different hyperparameters and data configurations. But when does this useful relationship break down? In this work, we explore its robustness. The key observation is that noisy data and the presence of nuisance features can be sufficient to shatter the Accuracy-on-the-line phenomenon. In these cases, ID and OOD accuracy can become negatively correlated, leading to "Accuracy-on-the-wrong-line". This phenomenon can also occur in the presence of spurious (shortcut) features, which tend to overshadow the more complex signal (core, non-spurious) features, resulting in a large nuisance feature space. Moreover, scaling to larger datasets does not mitigate this undesirable behavior and may even exacerbate it. We formally prove a lower bound on Out-of-distribution (OOD) error in a linear classification model, characterizing the conditions on the noise and nuisance features for a large OOD error. We finally demonstrate this phenomenon across both synthetic and real datasets with noisy data and nuisance features.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19049
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation
Sanyal, Amartya
Hu, Yaxi
Yu, Yaodong
Ma, Yian
Wang, Yixin
Schölkopf, Bernhard
Machine Learning
Artificial Intelligence
"Accuracy-on-the-line" is a widely observed phenomenon in machine learning, where a model's accuracy on in-distribution (ID) and out-of-distribution (OOD) data is positively correlated across different hyperparameters and data configurations. But when does this useful relationship break down? In this work, we explore its robustness. The key observation is that noisy data and the presence of nuisance features can be sufficient to shatter the Accuracy-on-the-line phenomenon. In these cases, ID and OOD accuracy can become negatively correlated, leading to "Accuracy-on-the-wrong-line". This phenomenon can also occur in the presence of spurious (shortcut) features, which tend to overshadow the more complex signal (core, non-spurious) features, resulting in a large nuisance feature space. Moreover, scaling to larger datasets does not mitigate this undesirable behavior and may even exacerbate it. We formally prove a lower bound on Out-of-distribution (OOD) error in a linear classification model, characterizing the conditions on the noise and nuisance features for a large OOD error. We finally demonstrate this phenomenon across both synthetic and real datasets with noisy data and nuisance features.
title Accuracy on the wrong line: On the pitfalls of noisy data for out-of-distribution generalisation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.19049