Multiple imputation and full law identifiability

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Karvanen, Juha, Tikka, Santtu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908617575235584
author Karvanen, Juha
Tikka, Santtu
author_facet Karvanen, Juha
Tikka, Santtu
contents The central challenges in missing data models concern the identifiability of two distributions: the target law and the full law. The target law refers to the joint distribution of the data variables, whereas the full law refers to the joint distribution of the data variables and their corresponding response indicators. However, the relationship between the identifiability of these two distributions and the feasibility of multiple imputation has not been clearly established. We show that imputations can be drawn from the correct conditional distributions for all possible missing data patterns if and only if the full law is identifiable. This result implies that standard multiple imputation methods -- which keep observed values unchanged and replace missing values with imputed values -- are invalid when the target law is identifiable but the full law is not. We demonstrate that alternative imputation strategies, in which certain observed values are also imputed, can enable the estimation of the target law in such cases.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18688
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multiple imputation and full law identifiability
Karvanen, Juha
Tikka, Santtu
Statistics Theory
62D10
The central challenges in missing data models concern the identifiability of two distributions: the target law and the full law. The target law refers to the joint distribution of the data variables, whereas the full law refers to the joint distribution of the data variables and their corresponding response indicators. However, the relationship between the identifiability of these two distributions and the feasibility of multiple imputation has not been clearly established. We show that imputations can be drawn from the correct conditional distributions for all possible missing data patterns if and only if the full law is identifiable. This result implies that standard multiple imputation methods -- which keep observed values unchanged and replace missing values with imputed values -- are invalid when the target law is identifiable but the full law is not. We demonstrate that alternative imputation strategies, in which certain observed values are also imputed, can enable the estimation of the target law in such cases.
title Multiple imputation and full law identifiability
topic Statistics Theory
62D10
url https://arxiv.org/abs/2410.18688