Statistical Challenges with Dataset Construction: Why You Will Never Have Enough Images

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Goldman, Josh, Tsotsos, John K.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929466649870336
author Goldman, Josh
Tsotsos, John K.
author_facet Goldman, Josh
Tsotsos, John K.
contents Deep neural networks have achieved impressive performance on many computer vision benchmarks in recent years. However, can we be confident that impressive performance on benchmarks will translate to strong performance in real-world environments? Many environments in the real world are safety critical, and even slight model failures can be catastrophic. Therefore, it is crucial to test models rigorously before deployment. We argue, through both statistical theory and empirical evidence, that selecting representative image datasets for testing a model is likely implausible in many domains. Furthermore, performance statistics calculated with non-representative image datasets are highly unreliable. As a consequence, we cannot guarantee that models which perform well on withheld test images will also perform well in the real world. Creating larger and larger datasets will not help, and bias aware datasets cannot solve this problem either. Ultimately, there is little statistical foundation for evaluating models using withheld test sets. We recommend that future evaluation methodologies focus on assessing a model's decision-making process, rather than metrics such as accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11160
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Statistical Challenges with Dataset Construction: Why You Will Never Have Enough Images
Goldman, Josh
Tsotsos, John K.
Computer Vision and Pattern Recognition
Computers and Society
Deep neural networks have achieved impressive performance on many computer vision benchmarks in recent years. However, can we be confident that impressive performance on benchmarks will translate to strong performance in real-world environments? Many environments in the real world are safety critical, and even slight model failures can be catastrophic. Therefore, it is crucial to test models rigorously before deployment. We argue, through both statistical theory and empirical evidence, that selecting representative image datasets for testing a model is likely implausible in many domains. Furthermore, performance statistics calculated with non-representative image datasets are highly unreliable. As a consequence, we cannot guarantee that models which perform well on withheld test images will also perform well in the real world. Creating larger and larger datasets will not help, and bias aware datasets cannot solve this problem either. Ultimately, there is little statistical foundation for evaluating models using withheld test sets. We recommend that future evaluation methodologies focus on assessing a model's decision-making process, rather than metrics such as accuracy.
title Statistical Challenges with Dataset Construction: Why You Will Never Have Enough Images
topic Computer Vision and Pattern Recognition
Computers and Society
url https://arxiv.org/abs/2408.11160