Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Morvan, Hugo, Agholme, Jonas, Eliasson, Bjorn, Olofsson, Katarina, Grote, Ludger, Iredahl, Fredrik, Sysoev, Oleg
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2601.21613
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908796741222400
author Morvan, Hugo
Agholme, Jonas
Eliasson, Bjorn
Olofsson, Katarina
Grote, Ludger
Iredahl, Fredrik
Sysoev, Oleg
author_facet Morvan, Hugo
Agholme, Jonas
Eliasson, Bjorn
Olofsson, Katarina
Grote, Ludger
Iredahl, Fredrik
Sysoev, Oleg
contents Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling multivariate missing data but traditional implementations face significant challenges when applied to big data sets due to computational time and memory limitations. To address this, the bigMICE package was developed, adapting the MICE framework to big data using Apache Spark MLLib and Spark ML. Our implementation allows for controlling the maximum memory usage during the execution, enabling processing of very large data sets on a hardware with a limited memory, such as ordinary laptops. The developed package was tested on a large Swedish medical registry to measure memory usage, runtime and dependence of the imputation quality on sample size and on missingness proportion in the data. In conclusion, our method is generally more memory efficient and faster on large data sets compared to a commonly used MICE implementation. We also demonstrate that working with very large datasets can result in high quality imputations even when a variable has a large proportion of missing data. This paper also provides guidelines and recommendations on how to install and use our open source package.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21613
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle bigMICE: Multiple Imputation of Big Data
Morvan, Hugo
Agholme, Jonas
Eliasson, Bjorn
Olofsson, Katarina
Grote, Ludger
Iredahl, Fredrik
Sysoev, Oleg
Computation
Distributed, Parallel, and Cluster Computing
Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling multivariate missing data but traditional implementations face significant challenges when applied to big data sets due to computational time and memory limitations. To address this, the bigMICE package was developed, adapting the MICE framework to big data using Apache Spark MLLib and Spark ML. Our implementation allows for controlling the maximum memory usage during the execution, enabling processing of very large data sets on a hardware with a limited memory, such as ordinary laptops. The developed package was tested on a large Swedish medical registry to measure memory usage, runtime and dependence of the imputation quality on sample size and on missingness proportion in the data. In conclusion, our method is generally more memory efficient and faster on large data sets compared to a commonly used MICE implementation. We also demonstrate that working with very large datasets can result in high quality imputations even when a variable has a large proportion of missing data. This paper also provides guidelines and recommendations on how to install and use our open source package.
title bigMICE: Multiple Imputation of Big Data
topic Computation
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2601.21613