Position: Measure Dataset Diversity, Don't Just Claim It

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Dora, Andrews, Jerone T. A., Papakyriakopoulos, Orestis, Xiang, Alice
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917719330258944
author Zhao, Dora
Andrews, Jerone T. A.
Papakyriakopoulos, Orestis
Xiang, Alice
author_facet Zhao, Dora
Andrews, Jerone T. A.
Papakyriakopoulos, Orestis
Xiang, Alice
contents Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets. Despite their prevalence, these terms lack clear definitions and validation. Our research explores the implications of this issue by analyzing "diversity" across 135 image and text datasets. Drawing from social sciences, we apply principles from measurement theory to identify considerations and offer recommendations for conceptualizing, operationalizing, and evaluating diversity in datasets. Our findings have broader implications for ML research, advocating for a more nuanced and precise approach to handling value-laden properties in dataset construction.
format Preprint
id arxiv_https___arxiv_org_abs_2407_08188
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Position: Measure Dataset Diversity, Don't Just Claim It
Zhao, Dora
Andrews, Jerone T. A.
Papakyriakopoulos, Orestis
Xiang, Alice
Machine Learning
Computers and Society
Machine learning (ML) datasets, often perceived as neutral, inherently encapsulate abstract and disputed social constructs. Dataset curators frequently employ value-laden terms such as diversity, bias, and quality to characterize datasets. Despite their prevalence, these terms lack clear definitions and validation. Our research explores the implications of this issue by analyzing "diversity" across 135 image and text datasets. Drawing from social sciences, we apply principles from measurement theory to identify considerations and offer recommendations for conceptualizing, operationalizing, and evaluating diversity in datasets. Our findings have broader implications for ML research, advocating for a more nuanced and precise approach to handling value-laden properties in dataset construction.
title Position: Measure Dataset Diversity, Don't Just Claim It
topic Machine Learning
Computers and Society
url https://arxiv.org/abs/2407.08188