GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramaswamy, Vikram V., Lin, Sing Yu, Zhao, Dora, Adcock, Aaron B., van der Maaten, Laurens, Ghadiyaram, Deepti, Russakovsky, Olga
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914032686989312
author Ramaswamy, Vikram V.
Lin, Sing Yu
Zhao, Dora
Adcock, Aaron B.
van der Maaten, Laurens
Ghadiyaram, Deepti
Russakovsky, Olga
author_facet Ramaswamy, Vikram V.
Lin, Sing Yu
Zhao, Dora
Adcock, Aaron B.
van der Maaten, Laurens
Ghadiyaram, Deepti
Russakovsky, Olga
contents Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, with no personally identifiable information, collected by soliciting images from people around the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. We demonstrate its use as both an evaluation and training dataset, allowing us to highlight and begin to mitigate the shortcomings in current models, despite GeoDE's relatively small size. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu
format Preprint
id arxiv_https___arxiv_org_abs_2301_02560
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
Ramaswamy, Vikram V.
Lin, Sing Yu
Zhao, Dora
Adcock, Aaron B.
van der Maaten, Laurens
Ghadiyaram, Deepti
Russakovsky, Olga
Computer Vision and Pattern Recognition
Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, with no personally identifiable information, collected by soliciting images from people around the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. We demonstrate its use as both an evaluation and training dataset, allowing us to highlight and begin to mitigate the shortcomings in current models, despite GeoDE's relatively small size. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu
title GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2301.02560