Croissant: A Metadata Format for ML-Ready Datasets
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866929620233748480 |
|---|---|
| author | Akhtar, Mubashara Benjelloun, Omar Conforti, Costanza Foschini, Luca Giner-Miguelez, Joan Gijsbers, Pieter Goswami, Sujata Jain, Nitisha Karamousadakis, Michalis Kuchnik, Michael Krishna, Satyapriya Lesage, Sylvain Lhoest, Quentin Marcenac, Pierre Maskey, Manil Mattson, Peter Oala, Luis Oderinwale, Hamidah Ruyssen, Pierre Santos, Tim Shinde, Rajat Simperl, Elena Suresh, Arjun Thomas, Goeffry Tykhonov, Slava Vanschoren, Joaquin Varma, Susheel van der Velde, Jos Vogler, Steffen Wu, Carole-Jean Zhang, Luyao |
| author_facet | Akhtar, Mubashara Benjelloun, Omar Conforti, Costanza Foschini, Luca Giner-Miguelez, Joan Gijsbers, Pieter Goswami, Sujata Jain, Nitisha Karamousadakis, Michalis Kuchnik, Michael Krishna, Satyapriya Lesage, Sylvain Lhoest, Quentin Marcenac, Pierre Maskey, Manil Mattson, Peter Oala, Luis Oderinwale, Hamidah Ruyssen, Pierre Santos, Tim Shinde, Rajat Simperl, Elena Suresh, Arjun Thomas, Goeffry Tykhonov, Slava Vanschoren, Joaquin Varma, Susheel van der Velde, Jos Vogler, Steffen Wu, Carole-Jean Zhang, Luyao |
| contents | Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_19546 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Croissant: A Metadata Format for ML-Ready Datasets Akhtar, Mubashara Benjelloun, Omar Conforti, Costanza Foschini, Luca Giner-Miguelez, Joan Gijsbers, Pieter Goswami, Sujata Jain, Nitisha Karamousadakis, Michalis Kuchnik, Michael Krishna, Satyapriya Lesage, Sylvain Lhoest, Quentin Marcenac, Pierre Maskey, Manil Mattson, Peter Oala, Luis Oderinwale, Hamidah Ruyssen, Pierre Santos, Tim Shinde, Rajat Simperl, Elena Suresh, Arjun Thomas, Goeffry Tykhonov, Slava Vanschoren, Joaquin Varma, Susheel van der Velde, Jos Vogler, Steffen Wu, Carole-Jean Zhang, Luyao Machine Learning Artificial Intelligence Databases Information Retrieval Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise. |
| title | Croissant: A Metadata Format for ML-Ready Datasets |
| topic | Machine Learning Artificial Intelligence Databases Information Retrieval |
| url | https://arxiv.org/abs/2403.19546 |