Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Salaudeen, Olawale, Zhang, Haoran, Alhamoud, Kumail, Beery, Sara, Ghassemi, Marzyeh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914121231892480
author Salaudeen, Olawale
Zhang, Haoran
Alhamoud, Kumail
Beery, Sara
Ghassemi, Marzyeh
author_facet Salaudeen, Olawale
Zhang, Haoran
Alhamoud, Kumail
Beery, Sara
Ghassemi, Marzyeh
contents Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-on-the-line." This pattern is often taken to imply that spurious correlations - correlations that improve ID but reduce OOD performance - are rare in practice. We find that this positive correlation is often an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy on the line does not hold. Across widely used distribution shift benchmarks, the OODSelect uncovers subsets, sometimes over half of the standard OOD set, where higher ID accuracy predicts lower OOD accuracy. Our findings indicate that aggregate metrics can obscure important failure modes of OOD robustness. We release code and the identified subsets to facilitate further research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24884
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Salaudeen, Olawale
Zhang, Haoran
Alhamoud, Kumail
Beery, Sara
Ghassemi, Marzyeh
Machine Learning
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-on-the-line." This pattern is often taken to imply that spurious correlations - correlations that improve ID but reduce OOD performance - are rare in practice. We find that this positive correlation is often an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy on the line does not hold. Across widely used distribution shift benchmarks, the OODSelect uncovers subsets, sometimes over half of the standard OOD set, where higher ID accuracy predicts lower OOD accuracy. Our findings indicate that aggregate metrics can obscure important failure modes of OOD robustness. We release code and the identified subsets to facilitate further research.
title Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
topic Machine Learning
url https://arxiv.org/abs/2510.24884