Multiple imputation in data that grow over time: A comparison of three strategies

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kavelaars, X. M., van Buuren, S., van Ginkel, J. R.
Format: Preprint
Publié: 2019
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916014116044800
author Kavelaars, X. M.
van Buuren, S.
van Ginkel, J. R.
author_facet Kavelaars, X. M.
van Buuren, S.
van Ginkel, J. R.
contents Multiple imputation is a highly recommended technique to deal with missing data, but the application to longitudinal datasets can be done in multiple ways. When a new wave of longitudinal data arrives, we can treat the combined data of multiple waves as a new missing data problem and overwrite existing imputations with new values (re-imputation). Alternatively, we may keep the existing imputations, and impute only the new data. We may do either a full multiple imputation (nested) or a single imputation (appended) on the new data per imputed set. This study compares these three strategies by means of simulation. All techniques resulted in valid inference under a monotone missingness pattern. A non-monotone missingness pattern led to biased and non-confidence valid regression coefficients after nested and appended imputation, depending on the correlation structure of the data. Correlations within timepoints must be stronger than correlations between timepoints to obtain valid inference. In an empirical example, the three strategies performed similarly.We conclude that appended imputation is especially beneficial in longitudinal datasets that suffer from dropout.
format Preprint
id arxiv_https___arxiv_org_abs_1904_04185
institution arXiv
publishDate 2019
record_format arxiv
spellingShingle Multiple imputation in data that grow over time: A comparison of three strategies
Kavelaars, X. M.
van Buuren, S.
van Ginkel, J. R.
Methodology
Multiple imputation is a highly recommended technique to deal with missing data, but the application to longitudinal datasets can be done in multiple ways. When a new wave of longitudinal data arrives, we can treat the combined data of multiple waves as a new missing data problem and overwrite existing imputations with new values (re-imputation). Alternatively, we may keep the existing imputations, and impute only the new data. We may do either a full multiple imputation (nested) or a single imputation (appended) on the new data per imputed set. This study compares these three strategies by means of simulation. All techniques resulted in valid inference under a monotone missingness pattern. A non-monotone missingness pattern led to biased and non-confidence valid regression coefficients after nested and appended imputation, depending on the correlation structure of the data. Correlations within timepoints must be stronger than correlations between timepoints to obtain valid inference. In an empirical example, the three strategies performed similarly.We conclude that appended imputation is especially beneficial in longitudinal datasets that suffer from dropout.
title Multiple imputation in data that grow over time: A comparison of three strategies
topic Methodology
url https://arxiv.org/abs/1904.04185