Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.21613 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908796741222400 |
|---|---|
| author | Morvan, Hugo Agholme, Jonas Eliasson, Bjorn Olofsson, Katarina Grote, Ludger Iredahl, Fredrik Sysoev, Oleg |
| author_facet | Morvan, Hugo Agholme, Jonas Eliasson, Bjorn Olofsson, Katarina Grote, Ludger Iredahl, Fredrik Sysoev, Oleg |
| contents | Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling multivariate missing data but traditional implementations face significant challenges when applied to big data sets due to computational time and memory limitations.
To address this, the bigMICE package was developed, adapting the MICE framework to big data using Apache Spark MLLib and Spark ML. Our implementation allows for controlling the maximum memory usage during the execution, enabling processing of very large data sets on a hardware with a limited memory, such as ordinary laptops.
The developed package was tested on a large Swedish medical registry to measure memory usage, runtime and dependence of the imputation quality on sample size and on missingness proportion in the data. In conclusion, our method is generally more memory efficient and faster on large data sets compared to a commonly used MICE implementation. We also demonstrate that working with very large datasets can result in high quality imputations even when a variable has a large proportion of missing data. This paper also provides guidelines and recommendations on how to install and use our open source package. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_21613 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | bigMICE: Multiple Imputation of Big Data Morvan, Hugo Agholme, Jonas Eliasson, Bjorn Olofsson, Katarina Grote, Ludger Iredahl, Fredrik Sysoev, Oleg Computation Distributed, Parallel, and Cluster Computing Missing data is a prevalent issue in many applications, including large medical registries such as the Swedish Healthcare Quality Registries, potentially leading to biased or inefficient analyses if not handled properly. Multiple Imputation by Chained Equations (MICE) is a popular and versatile method for handling multivariate missing data but traditional implementations face significant challenges when applied to big data sets due to computational time and memory limitations. To address this, the bigMICE package was developed, adapting the MICE framework to big data using Apache Spark MLLib and Spark ML. Our implementation allows for controlling the maximum memory usage during the execution, enabling processing of very large data sets on a hardware with a limited memory, such as ordinary laptops. The developed package was tested on a large Swedish medical registry to measure memory usage, runtime and dependence of the imputation quality on sample size and on missingness proportion in the data. In conclusion, our method is generally more memory efficient and faster on large data sets compared to a commonly used MICE implementation. We also demonstrate that working with very large datasets can result in high quality imputations even when a variable has a large proportion of missing data. This paper also provides guidelines and recommendations on how to install and use our open source package. |
| title | bigMICE: Multiple Imputation of Big Data |
| topic | Computation Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2601.21613 |