A Decade's Battle on Dataset Bias: Are We There Yet?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zhuang, He, Kaiming
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910852357029888
author Liu, Zhuang
He, Kaiming
author_facet Liu, Zhuang
He, Kaiming
contents We revisit the "dataset classification" experiment suggested by Torralba & Efros (2011) a decade ago, in the new era with large-scale, diverse, and hopefully less biased datasets as well as more capable neural network architectures. Surprisingly, we observe that modern neural networks can achieve excellent accuracy in classifying which dataset an image is from: e.g., we report 84.7% accuracy on held-out validation data for the three-way classification problem consisting of the YFCC, CC, and DataComp datasets. Our further experiments show that such a dataset classifier could learn semantic features that are generalizable and transferable, which cannot be explained by memorization. We hope our discovery will inspire the community to rethink issues involving dataset bias.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08632
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Decade's Battle on Dataset Bias: Are We There Yet?
Liu, Zhuang
He, Kaiming
Computer Vision and Pattern Recognition
Machine Learning
We revisit the "dataset classification" experiment suggested by Torralba & Efros (2011) a decade ago, in the new era with large-scale, diverse, and hopefully less biased datasets as well as more capable neural network architectures. Surprisingly, we observe that modern neural networks can achieve excellent accuracy in classifying which dataset an image is from: e.g., we report 84.7% accuracy on held-out validation data for the three-way classification problem consisting of the YFCC, CC, and DataComp datasets. Our further experiments show that such a dataset classifier could learn semantic features that are generalizable and transferable, which cannot be explained by memorization. We hope our discovery will inspire the community to rethink issues involving dataset bias.
title A Decade's Battle on Dataset Bias: Are We There Yet?
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2403.08632