Predicting Long Term Sequential Policy Value Using Softer Surrogates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nam, Hyunji, Nie, Allen, Gao, Ge, Syrgkanis, Vasilis, Brunskill, Emma
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913675294539776
author Nam, Hyunji
Nie, Allen
Gao, Ge
Syrgkanis, Vasilis
Brunskill, Emma
author_facet Nam, Hyunji
Nie, Allen
Gao, Ge
Syrgkanis, Vasilis
Brunskill, Emma
contents Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases when the new policy introduces novel actions. This issue commonly occurs in real-world domains, like healthcare, as new drugs and treatments are continuously developed. Novel actions necessitate on-policy data collection, which can be burdensome and expensive if the outcome of interest takes a substantial amount of time to observe--for example, in multi-year clinical trials. This raises a key question of how to predict the long-term outcome of a policy after only observing its short-term effects? Though in general this problem is intractable, under some surrogacy conditions, the short-term on-policy data can be combined with the long-term historical data to make accurate predictions about the new policy's long-term value. In two simulated healthcare examples--HIV and sepsis management--we show that our estimators can provide accurate predictions about the policy value only after observing 10\% of the full horizon data. We also provide finite sample analysis of our doubly robust estimators.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20638
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Predicting Long Term Sequential Policy Value Using Softer Surrogates
Nam, Hyunji
Nie, Allen
Gao, Ge
Syrgkanis, Vasilis
Brunskill, Emma
Artificial Intelligence
Machine Learning
Off-policy policy evaluation (OPE) estimates the outcome of a new policy using historical data collected from a different policy. However, existing OPE methods cannot handle cases when the new policy introduces novel actions. This issue commonly occurs in real-world domains, like healthcare, as new drugs and treatments are continuously developed. Novel actions necessitate on-policy data collection, which can be burdensome and expensive if the outcome of interest takes a substantial amount of time to observe--for example, in multi-year clinical trials. This raises a key question of how to predict the long-term outcome of a policy after only observing its short-term effects? Though in general this problem is intractable, under some surrogacy conditions, the short-term on-policy data can be combined with the long-term historical data to make accurate predictions about the new policy's long-term value. In two simulated healthcare examples--HIV and sepsis management--we show that our estimators can provide accurate predictions about the policy value only after observing 10\% of the full horizon data. We also provide finite sample analysis of our doubly robust estimators.
title Predicting Long Term Sequential Policy Value Using Softer Surrogates
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2412.20638