Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Kyungbok, Paik, Myunghee Cho
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929300294336512
author Lee, Kyungbok
Paik, Myunghee Cho
author_facet Lee, Kyungbok
Paik, Myunghee Cho
contents We introduce a novel doubly-robust (DR) off-policy evaluation (OPE) estimator for Markov decision processes, DRUnknown, designed for situations where both the logging policy and the value function are unknown. The proposed estimator initially estimates the logging policy and then estimates the value function model by minimizing the asymptotic variance of the estimator while considering the estimating effect of the logging policy. When the logging policy model is correctly specified, DRUnknown achieves the smallest asymptotic variance within the class containing existing OPE estimators. When the value function model is also correctly specified, DRUnknown is optimal as its asymptotic variance reaches the semiparametric lower bound. We present experimental results conducted in contextual bandits and reinforcement learning to compare the performance of DRUnknown with that of existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01830
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
Lee, Kyungbok
Paik, Myunghee Cho
Machine Learning
We introduce a novel doubly-robust (DR) off-policy evaluation (OPE) estimator for Markov decision processes, DRUnknown, designed for situations where both the logging policy and the value function are unknown. The proposed estimator initially estimates the logging policy and then estimates the value function model by minimizing the asymptotic variance of the estimator while considering the estimating effect of the logging policy. When the logging policy model is correctly specified, DRUnknown achieves the smallest asymptotic variance within the class containing existing OPE estimators. When the value function model is also correctly specified, DRUnknown is optimal as its asymptotic variance reaches the semiparametric lower bound. We present experimental results conducted in contextual bandits and reinforcement learning to compare the performance of DRUnknown with that of existing methods.
title Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
topic Machine Learning
url https://arxiv.org/abs/2404.01830