Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mohammed, Sedir, Naumann, Felix, Harmouch, Hazar
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912274704236544
author Mohammed, Sedir
Naumann, Felix
Harmouch, Hazar
author_facet Mohammed, Sedir
Naumann, Felix
Harmouch, Hazar
contents Data quality is crucial in machine learning (ML) applications, as errors in the data can significantly impact the prediction accuracy of the underlying ML model. Therefore, data cleaning is an integral component of any ML pipeline. However, in practical scenarios, data cleaning incurs significant costs, as it often involves domain experts for configuring and executing the cleaning process. Thus, efficient resource allocation during data cleaning can enhance ML prediction accuracy while controlling expenses. This paper presents COMET, a system designed to optimize data cleaning efforts for ML tasks. COMET gives step-by-step recommendations on which feature to clean next, maximizing the efficiency of data cleaning under resource constraints. We evaluated COMET across various datasets, ML algorithms, and data error types, demonstrating its robustness and adaptability. Our results show that COMET consistently outperforms feature importance-based, random, and another well-known cleaning method, achieving up to 52 and on average 5 percentage points higher ML prediction accuracy than the proposed baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11366
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy
Mohammed, Sedir
Naumann, Felix
Harmouch, Hazar
Databases
Data quality is crucial in machine learning (ML) applications, as errors in the data can significantly impact the prediction accuracy of the underlying ML model. Therefore, data cleaning is an integral component of any ML pipeline. However, in practical scenarios, data cleaning incurs significant costs, as it often involves domain experts for configuring and executing the cleaning process. Thus, efficient resource allocation during data cleaning can enhance ML prediction accuracy while controlling expenses. This paper presents COMET, a system designed to optimize data cleaning efforts for ML tasks. COMET gives step-by-step recommendations on which feature to clean next, maximizing the efficiency of data cleaning under resource constraints. We evaluated COMET across various datasets, ML algorithms, and data error types, demonstrating its robustness and adaptability. Our results show that COMET consistently outperforms feature importance-based, random, and another well-known cleaning method, achieving up to 52 and on average 5 percentage points higher ML prediction accuracy than the proposed baselines.
title Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy
topic Databases
url https://arxiv.org/abs/2503.11366