Logging Policy Design for Off-Policy Evaluation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Douglas, Connor, Persson, Joel, Provost, Foster
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910223375007744
author Douglas, Connor
Persson, Joel
Provost, Foster
author_facet Douglas, Connor
Persson, Joel
Provost, Foster
contents Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes experimentation without live deployment, yet in practice accuracy depends heavily on the logging policy used to collect data for computing the estimate. We study how to design logging policies that minimize OPE error for given target policies. We characterize a fundamental reward-coverage tradeoff: concentrating probability mass on high-reward actions reduces variance but risks missing signal on actions the target policy may take. We propose a unifying framework for logging policy design and derive optimal policies in canonical informational regimes where the target policy and reward distribution are (i) known, (ii) unknown, and (iii) partially known through priors or noisy estimates at logging time. Our results provide actionable guidance for firms choosing among multiple candidate recommendation systems. We demonstrate the importance of treatment selection when gathering data for OPE, and describe theoretically optimal approaches when this is a firm's primary objective. We also distill practical design principles for selecting logging policies when operational constraints prevent implementing the theoretical optimum.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15108
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Logging Policy Design for Off-Policy Evaluation
Douglas, Connor
Persson, Joel
Provost, Foster
Machine Learning
Artificial Intelligence
Information Retrieval
Methodology
62K05
G.3; I.2.6; I.2.8; H.3.3; H.2.8
Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes experimentation without live deployment, yet in practice accuracy depends heavily on the logging policy used to collect data for computing the estimate. We study how to design logging policies that minimize OPE error for given target policies. We characterize a fundamental reward-coverage tradeoff: concentrating probability mass on high-reward actions reduces variance but risks missing signal on actions the target policy may take. We propose a unifying framework for logging policy design and derive optimal policies in canonical informational regimes where the target policy and reward distribution are (i) known, (ii) unknown, and (iii) partially known through priors or noisy estimates at logging time. Our results provide actionable guidance for firms choosing among multiple candidate recommendation systems. We demonstrate the importance of treatment selection when gathering data for OPE, and describe theoretically optimal approaches when this is a firm's primary objective. We also distill practical design principles for selecting logging policies when operational constraints prevent implementing the theoretical optimum.
title Logging Policy Design for Off-Policy Evaluation
topic Machine Learning
Artificial Intelligence
Information Retrieval
Methodology
62K05
G.3; I.2.6; I.2.8; H.3.3; H.2.8
url https://arxiv.org/abs/2605.15108