Evaluating Model Performance Under Worst-case Subpopulations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Mike, Mittal, Daksh, Namkoong, Hongseok, Xia, Shangzhou
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914186752163840
author Li, Mike
Mittal, Daksh
Namkoong, Hongseok
Xia, Shangzhou
author_facet Li, Mike
Mittal, Daksh
Namkoong, Hongseok
Xia, Shangzhou
contents The performance of ML models degrades when the training population is different from that seen under operation. Towards assessing distributional robustness, we study the worst-case performance of a model over all subpopulations of a given size, defined with respect to core attributes Z. This notion of robustness can consider arbitrary (continuous) attributes Z, and automatically accounts for complex intersectionality in disadvantaged groups. We develop a scalable yet principled two-stage estimation procedure that can evaluate the robustness of state-of-the-art models. We prove that our procedure enjoys several finite-sample convergence guarantees, including dimension-free convergence. Instead of overly conservative notions based on Rademacher complexities, our evaluation error depends on the dimension of Z only through the out-of-sample error in estimating the performance conditional on Z. On real datasets, we demonstrate that our method certifies the robustness of a model and prevents deployment of unreliable models.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01316
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Model Performance Under Worst-case Subpopulations
Li, Mike
Mittal, Daksh
Namkoong, Hongseok
Xia, Shangzhou
Machine Learning
Computers and Society
The performance of ML models degrades when the training population is different from that seen under operation. Towards assessing distributional robustness, we study the worst-case performance of a model over all subpopulations of a given size, defined with respect to core attributes Z. This notion of robustness can consider arbitrary (continuous) attributes Z, and automatically accounts for complex intersectionality in disadvantaged groups. We develop a scalable yet principled two-stage estimation procedure that can evaluate the robustness of state-of-the-art models. We prove that our procedure enjoys several finite-sample convergence guarantees, including dimension-free convergence. Instead of overly conservative notions based on Rademacher complexities, our evaluation error depends on the dimension of Z only through the out-of-sample error in estimating the performance conditional on Z. On real datasets, we demonstrate that our method certifies the robustness of a model and prevents deployment of unreliable models.
title Evaluating Model Performance Under Worst-case Subpopulations
topic Machine Learning
Computers and Society
url https://arxiv.org/abs/2407.01316