Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahajan, Divyat, Mitliagkas, Ioannis, Neal, Brady, Syrgkanis, Vasilis
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917652615659520
author Mahajan, Divyat
Mitliagkas, Ioannis
Neal, Brady
Syrgkanis, Vasilis
author_facet Mahajan, Divyat
Mitliagkas, Ioannis
Neal, Brady
Syrgkanis, Vasilis
contents We study the problem of model selection in causal inference, specifically for conditional average treatment effect (CATE) estimation. Unlike machine learning, there is no perfect analogue of cross-validation for model selection as we do not observe the counterfactual potential outcomes. Towards this, a variety of surrogate metrics have been proposed for CATE model selection that use only observed data. However, we do not have a good understanding regarding their effectiveness due to limited comparisons in prior studies. We conduct an extensive empirical analysis to benchmark the surrogate model selection metrics introduced in the literature, as well as the novel ones introduced in this work. We ensure a fair comparison by tuning the hyperparameters associated with these metrics via AutoML, and provide more detailed trends by incorporating realistic datasets via generative modeling. Our analysis suggests novel model selection strategies based on careful hyperparameter selection of CATE estimators and causal ensembling.
format Preprint
id arxiv_https___arxiv_org_abs_2211_01939
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
Mahajan, Divyat
Mitliagkas, Ioannis
Neal, Brady
Syrgkanis, Vasilis
Machine Learning
Artificial Intelligence
Methodology
We study the problem of model selection in causal inference, specifically for conditional average treatment effect (CATE) estimation. Unlike machine learning, there is no perfect analogue of cross-validation for model selection as we do not observe the counterfactual potential outcomes. Towards this, a variety of surrogate metrics have been proposed for CATE model selection that use only observed data. However, we do not have a good understanding regarding their effectiveness due to limited comparisons in prior studies. We conduct an extensive empirical analysis to benchmark the surrogate model selection metrics introduced in the literature, as well as the novel ones introduced in this work. We ensure a fair comparison by tuning the hyperparameters associated with these metrics via AutoML, and provide more detailed trends by incorporating realistic datasets via generative modeling. Our analysis suggests novel model selection strategies based on careful hyperparameter selection of CATE estimators and causal ensembling.
title Empirical Analysis of Model Selection for Heterogeneous Causal Effect Estimation
topic Machine Learning
Artificial Intelligence
Methodology
url https://arxiv.org/abs/2211.01939