A Two-armed Bandit Framework for A/B Testing
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911074407677952 |
|---|---|
| author | Wang, Jinjuan Wen, Qianglin Zhang, Yu Yan, Xiaodong Shi, Chengchun |
| author_facet | Wang, Jinjuan Wen, Qianglin Zhang, Yu Yan, Xiaodong Shi, Chengchun |
| contents | A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_18118 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | A Two-armed Bandit Framework for A/B Testing Wang, Jinjuan Wen, Qianglin Zhang, Yu Yan, Xiaodong Shi, Chengchun Machine Learning Applications A/B testing is widely used in modern technology companies for policy evaluation and product deployment, with the goal of comparing the outcomes under a newly-developed policy against a standard control. Various causal inference and reinforcement learning methods developed in the literature are applicable to A/B testing. This paper introduces a two-armed bandit framework designed to improve the power of existing approaches. The proposed procedure consists of three main steps: (i) employing doubly robust estimation to generate pseudo-outcomes, (ii) utilizing a two-armed bandit framework to construct the test statistic, and (iii) applying a permutation-based method to compute the $p$-value. We demonstrate the efficacy of the proposed method through asymptotic theories, numerical experiments and real-world data from a ridesharing company, showing its superior performance in comparison to existing methods. |
| title | A Two-armed Bandit Framework for A/B Testing |
| topic | Machine Learning Applications |
| url | https://arxiv.org/abs/2507.18118 |