Multiple Imputation Methods under Extreme Values

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteur principal: Brasil, Enzo Porto
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911421906812928
author Brasil, Enzo Porto
author_facet Brasil, Enzo Porto
contents Missing data are ubiquitous in empirical databases, yet statistical analyses typically require complete data matrices. Multiple imputation offers a principled solution for filling these gaps. This study evaluates the performance of several multiple imputation methods, both in the presence and absence of extreme values, using the MICE package in R. Through Monte Carlo simulations, we generated incomplete data sets with three variables and assessed each imputation method within regression models. The results indicate that the linear regression based imputation method showed the best overall predictive performance (CV-MSE), whereas the sparse model approach was generally less efficient. Our findings underscore the relevance of extreme values when selecting an imputation strategy and highlight sample size, proportion of missingness, presence of extremes, and the type of fitted model as key determinants of performance. Despite its limitations, the study offers practical recommendations for researchers, stressing the need to examine the missingness mechanism and the occurrence of extreme values before choosing an imputation method.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04751
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multiple Imputation Methods under Extreme Values
Brasil, Enzo Porto
Computation
Methodology
62-08, 62F30, 62F35, 62G32, 62G05, 62H12, 62J20
Missing data are ubiquitous in empirical databases, yet statistical analyses typically require complete data matrices. Multiple imputation offers a principled solution for filling these gaps. This study evaluates the performance of several multiple imputation methods, both in the presence and absence of extreme values, using the MICE package in R. Through Monte Carlo simulations, we generated incomplete data sets with three variables and assessed each imputation method within regression models. The results indicate that the linear regression based imputation method showed the best overall predictive performance (CV-MSE), whereas the sparse model approach was generally less efficient. Our findings underscore the relevance of extreme values when selecting an imputation strategy and highlight sample size, proportion of missingness, presence of extremes, and the type of fitted model as key determinants of performance. Despite its limitations, the study offers practical recommendations for researchers, stressing the need to examine the missingness mechanism and the occurrence of extreme values before choosing an imputation method.
title Multiple Imputation Methods under Extreme Values
topic Computation
Methodology
62-08, 62F30, 62F35, 62G32, 62G05, 62H12, 62J20
url https://arxiv.org/abs/2602.04751