AI Competitions and Benchmarks: Dataset Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Egele, Romain, Junior, Julio C. S. Jacques, van Rijn, Jan N., Guyon, Isabelle, Baró, Xavier, Clapés, Albert, Balaprakash, Prasanna, Escalera, Sergio, Moeslund, Thomas, Wan, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916206392377344
author Egele, Romain
Junior, Julio C. S. Jacques
van Rijn, Jan N.
Guyon, Isabelle
Baró, Xavier
Clapés, Albert
Balaprakash, Prasanna
Escalera, Sergio
Moeslund, Thomas
Wan, Jun
author_facet Egele, Romain
Junior, Julio C. S. Jacques
van Rijn, Jan N.
Guyon, Isabelle
Baró, Xavier
Clapés, Albert
Balaprakash, Prasanna
Escalera, Sergio
Moeslund, Thomas
Wan, Jun
contents Machine learning is now used in many applications thanks to its ability to predict, generate, or discover patterns from large quantities of data. However, the process of collecting and transforming data for practical use is intricate. Even in today's digital era, where substantial data is generated daily, it is uncommon for it to be readily usable; most often, it necessitates meticulous manual data preparation. The haste in developing new models can frequently result in various shortcomings, potentially posing risks when deployed in real-world scenarios (eg social discrimination, critical failures), leading to the failure or substantial escalation of costs in AI-based projects. This chapter provides a comprehensive overview of established methodological tools, enriched by our practical experience, in the development of datasets for machine learning. Initially, we develop the tasks involved in dataset development and offer insights into their effective management (including requirements, design, implementation, evaluation, distribution, and maintenance). Then, we provide more details about the implementation process which includes data collection, transformation, and quality evaluation. Finally, we address practical considerations regarding dataset distribution and maintenance.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09703
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AI Competitions and Benchmarks: Dataset Development
Egele, Romain
Junior, Julio C. S. Jacques
van Rijn, Jan N.
Guyon, Isabelle
Baró, Xavier
Clapés, Albert
Balaprakash, Prasanna
Escalera, Sergio
Moeslund, Thomas
Wan, Jun
Machine Learning
Machine learning is now used in many applications thanks to its ability to predict, generate, or discover patterns from large quantities of data. However, the process of collecting and transforming data for practical use is intricate. Even in today's digital era, where substantial data is generated daily, it is uncommon for it to be readily usable; most often, it necessitates meticulous manual data preparation. The haste in developing new models can frequently result in various shortcomings, potentially posing risks when deployed in real-world scenarios (eg social discrimination, critical failures), leading to the failure or substantial escalation of costs in AI-based projects. This chapter provides a comprehensive overview of established methodological tools, enriched by our practical experience, in the development of datasets for machine learning. Initially, we develop the tasks involved in dataset development and offer insights into their effective management (including requirements, design, implementation, evaluation, distribution, and maintenance). Then, we provide more details about the implementation process which includes data collection, transformation, and quality evaluation. Finally, we address practical considerations regarding dataset distribution and maintenance.
title AI Competitions and Benchmarks: Dataset Development
topic Machine Learning
url https://arxiv.org/abs/2404.09703