To impute or not to impute: How machine learning modelers treat missing data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wanyi, Cummings, Mary
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912285586358272
author Chen, Wanyi
Cummings, Mary
author_facet Chen, Wanyi
Cummings, Mary
contents Missing data is prevalent in tabular machine learning (ML) models, and different missing data treatment methods can significantly affect ML model training results. However, little is known about how ML researchers and engineers choose missing data treatment methods and what factors affect their choices. To this end, we conducted a survey of 70 ML researchers and engineers. Our results revealed that most participants were not making informed decisions regarding missing data treatment, which could significantly affect the validity of the ML models trained by these researchers. We advocate for better education on missing data, more standardized missing data reporting, and better missing data analysis tools.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle To impute or not to impute: How machine learning modelers treat missing data
Chen, Wanyi
Cummings, Mary
Machine Learning
Human-Computer Interaction
Missing data is prevalent in tabular machine learning (ML) models, and different missing data treatment methods can significantly affect ML model training results. However, little is known about how ML researchers and engineers choose missing data treatment methods and what factors affect their choices. To this end, we conducted a survey of 70 ML researchers and engineers. Our results revealed that most participants were not making informed decisions regarding missing data treatment, which could significantly affect the validity of the ML models trained by these researchers. We advocate for better education on missing data, more standardized missing data reporting, and better missing data analysis tools.
title To impute or not to impute: How machine learning modelers treat missing data
topic Machine Learning
Human-Computer Interaction
url https://arxiv.org/abs/2503.16644