Croissant: A Metadata Format for ML-Ready Datasets

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Akhtar, Mubashara, Benjelloun, Omar, Conforti, Costanza, Foschini, Luca, Giner-Miguelez, Joan, Gijsbers, Pieter, Goswami, Sujata, Jain, Nitisha, Karamousadakis, Michalis, Kuchnik, Michael, Krishna, Satyapriya, Lesage, Sylvain, Lhoest, Quentin, Marcenac, Pierre, Maskey, Manil, Mattson, Peter, Oala, Luis, Oderinwale, Hamidah, Ruyssen, Pierre, Santos, Tim, Shinde, Rajat, Simperl, Elena, Suresh, Arjun, Thomas, Goeffry, Tykhonov, Slava, Vanschoren, Joaquin, Varma, Susheel, van der Velde, Jos, Vogler, Steffen, Wu, Carole-Jean, Zhang, Luyao
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929620233748480
author Akhtar, Mubashara
Benjelloun, Omar
Conforti, Costanza
Foschini, Luca
Giner-Miguelez, Joan
Gijsbers, Pieter
Goswami, Sujata
Jain, Nitisha
Karamousadakis, Michalis
Kuchnik, Michael
Krishna, Satyapriya
Lesage, Sylvain
Lhoest, Quentin
Marcenac, Pierre
Maskey, Manil
Mattson, Peter
Oala, Luis
Oderinwale, Hamidah
Ruyssen, Pierre
Santos, Tim
Shinde, Rajat
Simperl, Elena
Suresh, Arjun
Thomas, Goeffry
Tykhonov, Slava
Vanschoren, Joaquin
Varma, Susheel
van der Velde, Jos
Vogler, Steffen
Wu, Carole-Jean
Zhang, Luyao
author_facet Akhtar, Mubashara
Benjelloun, Omar
Conforti, Costanza
Foschini, Luca
Giner-Miguelez, Joan
Gijsbers, Pieter
Goswami, Sujata
Jain, Nitisha
Karamousadakis, Michalis
Kuchnik, Michael
Krishna, Satyapriya
Lesage, Sylvain
Lhoest, Quentin
Marcenac, Pierre
Maskey, Manil
Mattson, Peter
Oala, Luis
Oderinwale, Hamidah
Ruyssen, Pierre
Santos, Tim
Shinde, Rajat
Simperl, Elena
Suresh, Arjun
Thomas, Goeffry
Tykhonov, Slava
Vanschoren, Joaquin
Varma, Susheel
van der Velde, Jos
Vogler, Steffen
Wu, Carole-Jean
Zhang, Luyao
contents Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19546
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Croissant: A Metadata Format for ML-Ready Datasets
Akhtar, Mubashara
Benjelloun, Omar
Conforti, Costanza
Foschini, Luca
Giner-Miguelez, Joan
Gijsbers, Pieter
Goswami, Sujata
Jain, Nitisha
Karamousadakis, Michalis
Kuchnik, Michael
Krishna, Satyapriya
Lesage, Sylvain
Lhoest, Quentin
Marcenac, Pierre
Maskey, Manil
Mattson, Peter
Oala, Luis
Oderinwale, Hamidah
Ruyssen, Pierre
Santos, Tim
Shinde, Rajat
Simperl, Elena
Suresh, Arjun
Thomas, Goeffry
Tykhonov, Slava
Vanschoren, Joaquin
Varma, Susheel
van der Velde, Jos
Vogler, Steffen
Wu, Carole-Jean
Zhang, Luyao
Machine Learning
Artificial Intelligence
Databases
Information Retrieval
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise.
title Croissant: A Metadata Format for ML-Ready Datasets
topic Machine Learning
Artificial Intelligence
Databases
Information Retrieval
url https://arxiv.org/abs/2403.19546