Clustering Context in Off-Policy Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guzman-Olivares, Daniel, Schmidt, Philipp, Golebiowski, Jacek, Bekasov, Artur
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912252701966336
author Guzman-Olivares, Daniel
Schmidt, Philipp
Golebiowski, Jacek
Bekasov, Artur
author_facet Guzman-Olivares, Daniel
Schmidt, Philipp
Golebiowski, Jacek
Bekasov, Artur
contents Off-policy evaluation can leverage logged data to estimate the effectiveness of new policies in e-commerce, search engines, media streaming services, or automatic diagnostic tools in healthcare. However, the performance of baseline off-policy estimators like IPS deteriorates when the logging policy significantly differs from the evaluation policy. Recent work proposes sharing information across similar actions to mitigate this problem. In this work, we propose an alternative estimator that shares information across similar contexts using clustering. We study the theoretical properties of the proposed estimator, characterizing its bias and variance under different conditions. We also compare the performance of the proposed estimator and existing approaches in various synthetic problems, as well as a real-world recommendation dataset. Our experimental results confirm that clustering contexts improves estimation accuracy, especially in deficient information settings.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Clustering Context in Off-Policy Evaluation
Guzman-Olivares, Daniel
Schmidt, Philipp
Golebiowski, Jacek
Bekasov, Artur
Machine Learning
Artificial Intelligence
Off-policy evaluation can leverage logged data to estimate the effectiveness of new policies in e-commerce, search engines, media streaming services, or automatic diagnostic tools in healthcare. However, the performance of baseline off-policy estimators like IPS deteriorates when the logging policy significantly differs from the evaluation policy. Recent work proposes sharing information across similar actions to mitigate this problem. In this work, we propose an alternative estimator that shares information across similar contexts using clustering. We study the theoretical properties of the proposed estimator, characterizing its bias and variance under different conditions. We also compare the performance of the proposed estimator and existing approaches in various synthetic problems, as well as a real-world recommendation dataset. Our experimental results confirm that clustering contexts improves estimation accuracy, especially in deficient information settings.
title Clustering Context in Off-Policy Evaluation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2502.21304