RCT Rejection Sampling for Causal Estimation Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Keith, Katherine A., Feldman, Sergey, Jurgens, David, Bragg, Jonathan, Bhattacharya, Rohit
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910311736410112
author Keith, Katherine A.
Feldman, Sergey
Jurgens, David
Bragg, Jonathan
Bhattacharya, Rohit
author_facet Keith, Katherine A.
Feldman, Sergey
Jurgens, David
Bragg, Jonathan
Bhattacharya, Rohit
contents Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have proposed methods to adjust for confounding by adapting machine learning methods to the goal of causal estimation. However, empirical evaluation of these adjustment methods has been challenging and limited. In this work, we build on a promising empirical evaluation strategy that simplifies evaluation design and uses real data: subsampling randomized controlled trials (RCTs) to create confounded observational datasets while using the average causal effects from the RCTs as ground-truth. We contribute a new sampling algorithm, which we call RCT rejection sampling, and provide theoretical guarantees that causal identification holds in the observational data to allow for valid comparisons to the ground-truth RCT. Using synthetic data, we show our algorithm indeed results in low bias when oracle estimators are evaluated on the confounded samples, which is not always the case for a previously proposed algorithm. In addition to this identification result, we highlight several finite data considerations for evaluation designers who plan to use RCT rejection sampling on their own datasets. As a proof of concept, we implement an example evaluation pipeline and walk through these finite data considerations with a novel, real-world RCT -- which we release publicly -- consisting of approximately 70k observations and text data as high-dimensional covariates. Together, these contributions build towards a broader agenda of improved empirical evaluation for causal estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2307_15176
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RCT Rejection Sampling for Causal Estimation Evaluation
Keith, Katherine A.
Feldman, Sergey
Jurgens, David
Bragg, Jonathan
Bhattacharya, Rohit
Artificial Intelligence
Computation and Language
Machine Learning
Methodology
Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have proposed methods to adjust for confounding by adapting machine learning methods to the goal of causal estimation. However, empirical evaluation of these adjustment methods has been challenging and limited. In this work, we build on a promising empirical evaluation strategy that simplifies evaluation design and uses real data: subsampling randomized controlled trials (RCTs) to create confounded observational datasets while using the average causal effects from the RCTs as ground-truth. We contribute a new sampling algorithm, which we call RCT rejection sampling, and provide theoretical guarantees that causal identification holds in the observational data to allow for valid comparisons to the ground-truth RCT. Using synthetic data, we show our algorithm indeed results in low bias when oracle estimators are evaluated on the confounded samples, which is not always the case for a previously proposed algorithm. In addition to this identification result, we highlight several finite data considerations for evaluation designers who plan to use RCT rejection sampling on their own datasets. As a proof of concept, we implement an example evaluation pipeline and walk through these finite data considerations with a novel, real-world RCT -- which we release publicly -- consisting of approximately 70k observations and text data as high-dimensional covariates. Together, these contributions build towards a broader agenda of improved empirical evaluation for causal estimation.
title RCT Rejection Sampling for Causal Estimation Evaluation
topic Artificial Intelligence
Computation and Language
Machine Learning
Methodology
url https://arxiv.org/abs/2307.15176