Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bohacek, Matyas, Echavarri, Ignacio Vilanova
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915694226964480
author Bohacek, Matyas
Echavarri, Ignacio Vilanova
author_facet Bohacek, Matyas
Echavarri, Ignacio Vilanova
contents Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets are often built using unrestricted and opaque data collection practices. While most literature focuses on the development and applications of GAI models, the ethical and legal considerations surrounding the creation of these datasets are often neglected. In addition, as datasets are shared, edited, and further reproduced online, information about their origin, legitimacy, and safety often gets lost. To address this gap, we introduce the Compliance Rating Scheme (CRS), a framework designed to evaluate dataset compliance with critical transparency, accountability, and security principles. We also release an open-source Python library built around data provenance technology to implement this framework, allowing for seamless integration into existing dataset-processing and AI training pipelines. The library is simultaneously reactive and proactive, as in addition to evaluating the CRS of existing datasets, it equally informs responsible scraping and construction of new datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21775
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
Bohacek, Matyas
Echavarri, Ignacio Vilanova
Artificial Intelligence
Computers and Society
Databases
Generative Artificial Intelligence (GAI) has experienced exponential growth in recent years, partly facilitated by the abundance of large-scale open-source datasets. These datasets are often built using unrestricted and opaque data collection practices. While most literature focuses on the development and applications of GAI models, the ethical and legal considerations surrounding the creation of these datasets are often neglected. In addition, as datasets are shared, edited, and further reproduced online, information about their origin, legitimacy, and safety often gets lost. To address this gap, we introduce the Compliance Rating Scheme (CRS), a framework designed to evaluate dataset compliance with critical transparency, accountability, and security principles. We also release an open-source Python library built around data provenance technology to implement this framework, allowing for seamless integration into existing dataset-processing and AI training pipelines. The library is simultaneously reactive and proactive, as in addition to evaluating the CRS of existing datasets, it equally informs responsible scraping and construction of new datasets.
title Compliance Rating Scheme: A Data Provenance Framework for Generative AI Datasets
topic Artificial Intelligence
Computers and Society
Databases
url https://arxiv.org/abs/2512.21775