Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Winata, Genta Indra, Anugraha, David, Liu, Emmy, Aji, Alham Fikri, Hung, Shou-Yi, Parashar, Aditya, Irawan, Patrick Amadeus, Zhang, Ruochen, Yong, Zheng-Xin, Cruz, Jan Christian Blaise, Muennighoff, Niklas, Kim, Seungone, Zhao, Hanyang, Kar, Sudipta, Suryoraharjo, Kezia Erina, Adilazuarda, M. Farid, Lee, En-Shiun Annie, Purwarianti, Ayu, Wijaya, Derry Tanti, Choudhury, Monojit
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912410748583936
author Winata, Genta Indra
Anugraha, David
Liu, Emmy
Aji, Alham Fikri
Hung, Shou-Yi
Parashar, Aditya
Irawan, Patrick Amadeus
Zhang, Ruochen
Yong, Zheng-Xin
Cruz, Jan Christian Blaise
Muennighoff, Niklas
Kim, Seungone
Zhao, Hanyang
Kar, Sudipta
Suryoraharjo, Kezia Erina
Adilazuarda, M. Farid
Lee, En-Shiun Annie
Purwarianti, Ayu
Wijaya, Derry Tanti
Choudhury, Monojit
author_facet Winata, Genta Indra
Anugraha, David
Liu, Emmy
Aji, Alham Fikri
Hung, Shou-Yi
Parashar, Aditya
Irawan, Patrick Amadeus
Zhang, Ruochen
Yong, Zheng-Xin
Cruz, Jan Christian Blaise
Muennighoff, Niklas
Kim, Seungone
Zhao, Hanyang
Kar, Sudipta
Suryoraharjo, Kezia Erina
Adilazuarda, M. Farid
Lee, En-Shiun Annie
Purwarianti, Ayu
Wijaya, Derry Tanti
Choudhury, Monojit
contents High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challenge. Many dataset paper submissions lack originality, diversity, or rigorous quality control, and these shortcomings are often overlooked during peer review. Submissions also frequently omit essential details about dataset construction and properties. While existing tools such as datasheets aim to promote transparency, they are largely descriptive and do not provide standardized, measurable methods for evaluating data quality. Similarly, metadata requirements at conferences promote accountability but are inconsistently enforced. To address these limitations, this position paper advocates for the integration of systematic, rubric-based evaluation metrics into the dataset review process-particularly as submission volumes continue to grow. We also explore scalable, cost-effective methods for synthetic data generation, including dedicated tools and LLM-as-a-judge approaches, to support more efficient evaluation. As a call to action, we introduce DataRubrics, a structured framework for assessing the quality of both human- and model-generated datasets. Leveraging recent advances in LLM-based evaluation, DataRubrics offers a reproducible, scalable, and actionable solution for dataset quality assessment, enabling both authors and reviewers to uphold higher standards in data-centric research. We also release code to support reproducibility of LLM-based evaluations at https://github.com/datarubrics/datarubrics.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
Winata, Genta Indra
Anugraha, David
Liu, Emmy
Aji, Alham Fikri
Hung, Shou-Yi
Parashar, Aditya
Irawan, Patrick Amadeus
Zhang, Ruochen
Yong, Zheng-Xin
Cruz, Jan Christian Blaise
Muennighoff, Niklas
Kim, Seungone
Zhao, Hanyang
Kar, Sudipta
Suryoraharjo, Kezia Erina
Adilazuarda, M. Farid
Lee, En-Shiun Annie
Purwarianti, Ayu
Wijaya, Derry Tanti
Choudhury, Monojit
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Audio and Speech Processing
High-quality datasets are fundamental to training and evaluating machine learning models, yet their creation-especially with accurate human annotations-remains a significant challenge. Many dataset paper submissions lack originality, diversity, or rigorous quality control, and these shortcomings are often overlooked during peer review. Submissions also frequently omit essential details about dataset construction and properties. While existing tools such as datasheets aim to promote transparency, they are largely descriptive and do not provide standardized, measurable methods for evaluating data quality. Similarly, metadata requirements at conferences promote accountability but are inconsistently enforced. To address these limitations, this position paper advocates for the integration of systematic, rubric-based evaluation metrics into the dataset review process-particularly as submission volumes continue to grow. We also explore scalable, cost-effective methods for synthetic data generation, including dedicated tools and LLM-as-a-judge approaches, to support more efficient evaluation. As a call to action, we introduce DataRubrics, a structured framework for assessing the quality of both human- and model-generated datasets. Leveraging recent advances in LLM-based evaluation, DataRubrics offers a reproducible, scalable, and actionable solution for dataset quality assessment, enabling both authors and reviewers to uphold higher standards in data-centric research. We also release code to support reproducibility of LLM-based evaluations at https://github.com/datarubrics/datarubrics.
title Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2506.01789