Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Xiangkun, Wen, Qianglin, Zhang, Yingying, Zhu, Hongtu, Li, Ting, Shi, Chengchun
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910008621400064
author Wu, Xiangkun
Wen, Qianglin
Zhang, Yingying
Zhu, Hongtu
Li, Ting
Shi, Chengchun
author_facet Wu, Xiangkun
Wen, Qianglin
Zhang, Yingying
Zhu, Hongtu
Li, Ting
Shi, Chengchun
contents A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning
Wu, Xiangkun
Wen, Qianglin
Zhang, Yingying
Zhu, Hongtu
Li, Ting
Shi, Chengchun
Machine Learning
Methodology
A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where policies are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.
title Designing Time Series Experiments in A/B Testing with Transformer Reinforcement Learning
topic Machine Learning
Methodology
url https://arxiv.org/abs/2602.01853