CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mandyam, Aishwarya, Tang, Shengpu, Yao, Jiayu, Wiens, Jenna, Engelhardt, Barbara E.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913164761759744
author Mandyam, Aishwarya
Tang, Shengpu
Yao, Jiayu
Wiens, Jenna
Engelhardt, Barbara E.
author_facet Mandyam, Aishwarya
Tang, Shengpu
Yao, Jiayu
Wiens, Jenna
Engelhardt, Barbara E.
contents Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be evaluated prior to deployment. Unfortunately, OPE techniques are inherently limited by the breadth of the available data, which may not be sufficient to evaluate the performance of a new policy. Recent work attempts to improve dataset coverage by adding expert-annotated counterfactual samples. However, such annotations are often imperfect and can lead to worse estimator performance than using no annotations at all. To better leverage imperfect annotations, we propose a family of OPE estimators grounded in the doubly robust (DR) framework, which combines importance sampling (IS) with a reward model (direct method, DM) for better statistical guarantees. We study three ways of incorporating counterfactual annotations. Under mild assumptions, we prove that using annotations within just the DM component yields the most desirable theoretical results. Experiments on multiple healthcare tasks, including real-world electronic health records (EHR) data, show that this strategy is most robust under misspecified reward models and inaccurate annotations. By addressing the challenges posed by imperfect annotations, this work broadens the applicability of OPE methods and facilitates safer deployment of decision-making policies in healthcare.
format Preprint
id arxiv_https___arxiv_org_abs_2412_08052
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
Mandyam, Aishwarya
Tang, Shengpu
Yao, Jiayu
Wiens, Jenna
Engelhardt, Barbara E.
Machine Learning
Off-policy evaluation (OPE) is critical for applying contextual bandit algorithms to high-stakes decision-making settings such as healthcare, where new treatment policies must be evaluated prior to deployment. Unfortunately, OPE techniques are inherently limited by the breadth of the available data, which may not be sufficient to evaluate the performance of a new policy. Recent work attempts to improve dataset coverage by adding expert-annotated counterfactual samples. However, such annotations are often imperfect and can lead to worse estimator performance than using no annotations at all. To better leverage imperfect annotations, we propose a family of OPE estimators grounded in the doubly robust (DR) framework, which combines importance sampling (IS) with a reward model (direct method, DM) for better statistical guarantees. We study three ways of incorporating counterfactual annotations. Under mild assumptions, we prove that using annotations within just the DM component yields the most desirable theoretical results. Experiments on multiple healthcare tasks, including real-world electronic health records (EHR) data, show that this strategy is most robust under misspecified reward models and inaccurate annotations. By addressing the challenges posed by imperfect annotations, this work broadens the applicability of OPE methods and facilitates safer deployment of decision-making policies in healthcare.
title CANDOR: Counterfactual ANnotated DOubly Robust Off-Policy Evaluation
topic Machine Learning
url https://arxiv.org/abs/2412.08052