A Two-armed Bandit Framework for A/B Testing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jinjuan, Wen, Qianglin, Zhang, Yu, Yan, Xiaodong, Shi, Chengchun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911074407677952
author Wang, Jinjuan
Wen, Qianglin
Zhang, Yu
Yan, Xiaodong
Shi, Chengchun
author_facet Wang, Jinjuan
Wen, Qianglin
Zhang, Yu
Yan, Xiaodong
Shi, Chengchun
contents A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Two-armed Bandit Framework for A/B Testing
Wang, Jinjuan
Wen, Qianglin
Zhang, Yu
Yan, Xiaodong
Shi, Chengchun
Machine Learning
Applications
A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods.
title A Two-armed Bandit Framework for A/B Testing
topic Machine Learning
Applications
url https://arxiv.org/abs/2507.18118