We Need Improved Data Curation and Attribution in AI for Scientific Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Graziani, Mara, Foncubierta, Antonio, Christofidellis, Dimitrios, Espejo-Morales, Irina, Molnar, Malina, Alberts, Marvin, Manica, Matteo, Born, Jannis
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912307314950144
author Graziani, Mara
Foncubierta, Antonio
Christofidellis, Dimitrios
Espejo-Morales, Irina
Molnar, Malina
Alberts, Marvin
Manica, Matteo
Born, Jannis
author_facet Graziani, Mara
Foncubierta, Antonio
Christofidellis, Dimitrios
Espejo-Morales, Irina
Molnar, Malina
Alberts, Marvin
Manica, Matteo
Born, Jannis
contents As the interplay between human-generated and synthetic data evolves, new challenges arise in scientific discovery concerning the integrity of the data and the stability of the models. In this work, we examine the role of synthetic data as opposed to that of real experimental data for scientific research. Our analyses indicate that nearly three-quarters of experimental datasets available on open-access platforms have relatively low adoption rates, opening new opportunities to enhance their discoverability and usability by automated methods. Additionally, we observe an increasing difficulty in distinguishing synthetic from real experimental data. We propose supplementing ongoing efforts in automating synthetic data detection by increasing the focus on watermarking real experimental data, thereby strengthening data traceability and integrity. Our estimates suggest that watermarking even less than half of the real world data generated annually could help sustain model robustness, while promoting a balanced integration of synthetic and human-generated content.
format Preprint
id arxiv_https___arxiv_org_abs_2504_02486
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle We Need Improved Data Curation and Attribution in AI for Scientific Discovery
Graziani, Mara
Foncubierta, Antonio
Christofidellis, Dimitrios
Espejo-Morales, Irina
Molnar, Malina
Alberts, Marvin
Manica, Matteo
Born, Jannis
Artificial Intelligence
As the interplay between human-generated and synthetic data evolves, new challenges arise in scientific discovery concerning the integrity of the data and the stability of the models. In this work, we examine the role of synthetic data as opposed to that of real experimental data for scientific research. Our analyses indicate that nearly three-quarters of experimental datasets available on open-access platforms have relatively low adoption rates, opening new opportunities to enhance their discoverability and usability by automated methods. Additionally, we observe an increasing difficulty in distinguishing synthetic from real experimental data. We propose supplementing ongoing efforts in automating synthetic data detection by increasing the focus on watermarking real experimental data, thereby strengthening data traceability and integrity. Our estimates suggest that watermarking even less than half of the real world data generated annually could help sustain model robustness, while promoting a balanced integration of synthetic and human-generated content.
title We Need Improved Data Curation and Attribution in AI for Scientific Discovery
topic Artificial Intelligence
url https://arxiv.org/abs/2504.02486