Offline Contextual Bandit with Counterfactual Sample Identification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gilotte, Alexandre, Sakhi, Otmane, Aouali, Imad, Heymann, Benjamin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911152547561472
author Gilotte, Alexandre
Sakhi, Otmane
Aouali, Imad
Heymann, Benjamin
author_facet Gilotte, Alexandre
Sakhi, Otmane
Aouali, Imad
Heymann, Benjamin
contents In production systems, contextual bandit approaches often rely on direct reward models that take both action and context as input. However, these models can suffer from confounding, making it difficult to isolate the effect of the action from that of the context. We present \emph{Counterfactual Sample Identification}, a new approach that re-frames the problem: rather than predicting reward, it learns to recognize which action led to a successful (binary) outcome by comparing it to a counterfactual action sampled from the logging policy under the same context. The method is theoretically grounded and consistently outperforms direct models in both synthetic experiments and real-world deployments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Offline Contextual Bandit with Counterfactual Sample Identification
Gilotte, Alexandre
Sakhi, Otmane
Aouali, Imad
Heymann, Benjamin
Machine Learning
In production systems, contextual bandit approaches often rely on direct reward models that take both action and context as input. However, these models can suffer from confounding, making it difficult to isolate the effect of the action from that of the context. We present \emph{Counterfactual Sample Identification}, a new approach that re-frames the problem: rather than predicting reward, it learns to recognize which action led to a successful (binary) outcome by comparing it to a counterfactual action sampled from the logging policy under the same context. The method is theoretically grounded and consistently outperforms direct models in both synthetic experiments and real-world deployments.
title Offline Contextual Bandit with Counterfactual Sample Identification
topic Machine Learning
url https://arxiv.org/abs/2509.10520