GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914032686989312 |
|---|---|
| author | Ramaswamy, Vikram V. Lin, Sing Yu Zhao, Dora Adcock, Aaron B. van der Maaten, Laurens Ghadiyaram, Deepti Russakovsky, Olga |
| author_facet | Ramaswamy, Vikram V. Lin, Sing Yu Zhao, Dora Adcock, Aaron B. van der Maaten, Laurens Ghadiyaram, Deepti Russakovsky, Olga |
| contents | Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, with no personally identifiable information, collected by soliciting images from people around the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. We demonstrate its use as both an evaluation and training dataset, allowing us to highlight and begin to mitigate the shortcomings in current models, despite GeoDE's relatively small size. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2301_02560 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition Ramaswamy, Vikram V. Lin, Sing Yu Zhao, Dora Adcock, Aaron B. van der Maaten, Laurens Ghadiyaram, Deepti Russakovsky, Olga Computer Vision and Pattern Recognition Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce stereotypical biases, can contain personally identifiable information, and typically originates from Europe and North America. In this work, we rethink the dataset collection paradigm and introduce GeoDE, a geographically diverse dataset with 61,940 images from 40 classes and 6 world regions, with no personally identifiable information, collected by soliciting images from people around the world. We analyse GeoDE to understand differences in images collected in this manner compared to web-scraping. We demonstrate its use as both an evaluation and training dataset, allowing us to highlight and begin to mitigate the shortcomings in current models, despite GeoDE's relatively small size. We release the full dataset and code at https://geodiverse-data-collection.cs.princeton.edu |
| title | GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2301.02560 |