Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Rondina, Marco, Vetrò, Antonio, De Martin, Juan Carlos
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909540722671616
author Rondina, Marco
Vetrò, Antonio
De Martin, Juan Carlos
author_facet Rondina, Marco
Vetrò, Antonio
De Martin, Juan Carlos
contents ML/AI is the field of computer science and computer engineering that arguably received the most attention and funding over the last decade. Data is the key element of ML/AI, so it is becoming increasingly important to ensure that users are fully aware of the quality of the datasets that they use, and of the process generating them, so that possible negative impacts on downstream effects can be tracked, analysed, and, where possible, mitigated. One of the tools that can be useful in this perspective is dataset documentation. The aim of this work is to investigate the state of dataset documentation practices, measuring the completeness of the documentation of several popular datasets in ML/AI repositories. We created a dataset documentation schema -- the Documentation Test Sheet (DTS) -- that identifies the information that should always be attached to a dataset (to ensure proper dataset choice and informed use), according to relevant studies in the literature. We verified 100 popular datasets from four different repositories with the DTS to investigate which information was present. Overall, we observed a lack of relevant documentation, especially about the context of data collection and data processing, highlighting a paucity of transparency.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13463
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
Rondina, Marco
Vetrò, Antonio
De Martin, Juan Carlos
Digital Libraries
Artificial Intelligence
Human-Computer Interaction
Machine Learning
ML/AI is the field of computer science and computer engineering that arguably received the most attention and funding over the last decade. Data is the key element of ML/AI, so it is becoming increasingly important to ensure that users are fully aware of the quality of the datasets that they use, and of the process generating them, so that possible negative impacts on downstream effects can be tracked, analysed, and, where possible, mitigated. One of the tools that can be useful in this perspective is dataset documentation. The aim of this work is to investigate the state of dataset documentation practices, measuring the completeness of the documentation of several popular datasets in ML/AI repositories. We created a dataset documentation schema -- the Documentation Test Sheet (DTS) -- that identifies the information that should always be attached to a dataset (to ensure proper dataset choice and informed use), according to relevant studies in the literature. We verified 100 popular datasets from four different repositories with the DTS to investigate which information was present. Overall, we observed a lack of relevant documentation, especially about the context of data collection and data processing, highlighting a paucity of transparency.
title Completeness of Datasets Documentation on ML/AI repositories: an Empirical Investigation
topic Digital Libraries
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2503.13463