Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Testa, Lorenzo, Kennedy, Edward H., Reimherr, Matthew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911394990915584
author Testa, Lorenzo
Kennedy, Edward H.
Reimherr, Matthew
author_facet Testa, Lorenzo
Kennedy, Edward H.
Reimherr, Matthew
contents The Difference-in-Differences (DiD) method is a fundamental tool for causal inference, yet its application is often complicated by missing data. Although recent work has developed robust DiD estimators for complex settings like staggered treatment adoption, these methods typically assume complete data and fail to address the critical challenge of outcomes that are missing at random (MAR) -- a common problem that invalidates standard estimators. We develop a rigorous framework, rooted in semiparametric theory, for identifying and efficiently estimating the Average Treatment Effect on the Treated (ATT) when either pre- or post-treatment (or both) outcomes are missing at random. We first establish nonparametric identification of the ATT under two minimal sets of sufficient conditions. For each, we derive the semiparametric efficiency bound, which provides a formal benchmark for asymptotic optimality. We then propose novel estimators that are asymptotically efficient, achieving this theoretical bound. A key feature of our estimators is their multiple robustness, which ensures consistency even if some nuisance function models are misspecified. We validate the properties of our estimators and showcase their broad applicability through an extensive simulation study.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25009
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random
Testa, Lorenzo
Kennedy, Edward H.
Reimherr, Matthew
Methodology
Econometrics
Statistics Theory
The Difference-in-Differences (DiD) method is a fundamental tool for causal inference, yet its application is often complicated by missing data. Although recent work has developed robust DiD estimators for complex settings like staggered treatment adoption, these methods typically assume complete data and fail to address the critical challenge of outcomes that are missing at random (MAR) -- a common problem that invalidates standard estimators. We develop a rigorous framework, rooted in semiparametric theory, for identifying and efficiently estimating the Average Treatment Effect on the Treated (ATT) when either pre- or post-treatment (or both) outcomes are missing at random. We first establish nonparametric identification of the ATT under two minimal sets of sufficient conditions. For each, we derive the semiparametric efficiency bound, which provides a formal benchmark for asymptotic optimality. We then propose novel estimators that are asymptotically efficient, achieving this theoretical bound. A key feature of our estimators is their multiple robustness, which ensures consistency even if some nuisance function models are misspecified. We validate the properties of our estimators and showcase their broad applicability through an extensive simulation study.
title Efficient Difference-in-Differences Estimation when Outcomes are Missing at Random
topic Methodology
Econometrics
Statistics Theory
url https://arxiv.org/abs/2509.25009