Robustness to Subpopulation Shift with Domain Label Noise via Regularized Annotation of Domains

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Stromberg, Nathan, Ayyagari, Rohan, Welfert, Monica, Koyejo, Sanmi, Nock, Richard, Sankar, Lalitha
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913405140467712
author Stromberg, Nathan
Ayyagari, Rohan
Welfert, Monica
Koyejo, Sanmi
Nock, Richard
Sankar, Lalitha
author_facet Stromberg, Nathan
Ayyagari, Rohan
Welfert, Monica
Koyejo, Sanmi
Nock, Richard
Sankar, Lalitha
contents Existing methods for last layer retraining that aim to optimize worst-group accuracy (WGA) rely heavily on well-annotated groups in the training data. We show, both in theory and practice, that annotation-based data augmentations using either downsampling or upweighting for WGA are susceptible to domain annotation noise, and in high-noise regimes approach the WGA of a model trained with vanilla empirical risk minimization. We introduce Regularized Annotation of Domains (RAD) in order to train robust last layer classifiers without the need for explicit domain annotations. Our results show that RAD is competitive with other recently proposed domain annotation-free techniques. Most importantly, RAD outperforms state-of-the-art annotation-reliant methods even with only 5% noise in the training data for several publicly available datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11039
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robustness to Subpopulation Shift with Domain Label Noise via Regularized Annotation of Domains
Stromberg, Nathan
Ayyagari, Rohan
Welfert, Monica
Koyejo, Sanmi
Nock, Richard
Sankar, Lalitha
Machine Learning
Existing methods for last layer retraining that aim to optimize worst-group accuracy (WGA) rely heavily on well-annotated groups in the training data. We show, both in theory and practice, that annotation-based data augmentations using either downsampling or upweighting for WGA are susceptible to domain annotation noise, and in high-noise regimes approach the WGA of a model trained with vanilla empirical risk minimization. We introduce Regularized Annotation of Domains (RAD) in order to train robust last layer classifiers without the need for explicit domain annotations. Our results show that RAD is competitive with other recently proposed domain annotation-free techniques. Most importantly, RAD outperforms state-of-the-art annotation-reliant methods even with only 5% noise in the training data for several publicly available datasets.
title Robustness to Subpopulation Shift with Domain Label Noise via Regularized Annotation of Domains
topic Machine Learning
url https://arxiv.org/abs/2402.11039