Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Johari, Ramesh, Peng, Tianyi, Xing, Wenqian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915536373284864
author Johari, Ramesh
Peng, Tianyi
Xing, Wenqian
author_facet Johari, Ramesh
Peng, Tianyi
Xing, Wenqian
contents Randomized experiments (or A/B tests) are widely used to evaluate interventions in dynamic systems such as recommendation platforms, marketplaces, and digital health. In these settings, interventions affect both current and future system states, so estimating the global average treatment effect (GATE) requires accounting for temporal dynamics, which is especially challenging in the presence of nonstationarity; existing approaches suffer from high bias, high variance, or both. In this paper, we address this challenge via the novel Truncated Policy Gradient (TPG) estimator, which replaces instantaneous outcomes with short-horizon outcome trajectories. The estimator admits a policy-gradient interpretation: it is a truncation of the first-order approximation to the GATE, yielding provable reductions in bias and variance in nonstationary Markovian settings. We further establish a central limit theorem for the TPG estimator and develop a consistent variance estimator that remains valid under nonstationarity with single-trajectory data. We validate our theory with two real-world case studies. The results show that a well-calibrated TPG estimator attains low bias and variance in practical nonstationary settings.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator
Johari, Ramesh
Peng, Tianyi
Xing, Wenqian
Methodology
Randomized experiments (or A/B tests) are widely used to evaluate interventions in dynamic systems such as recommendation platforms, marketplaces, and digital health. In these settings, interventions affect both current and future system states, so estimating the global average treatment effect (GATE) requires accounting for temporal dynamics, which is especially challenging in the presence of nonstationarity; existing approaches suffer from high bias, high variance, or both. In this paper, we address this challenge via the novel Truncated Policy Gradient (TPG) estimator, which replaces instantaneous outcomes with short-horizon outcome trajectories. The estimator admits a policy-gradient interpretation: it is a truncation of the first-order approximation to the GATE, yielding provable reductions in bias and variance in nonstationary Markovian settings. We further establish a central limit theorem for the TPG estimator and develop a consistent variance estimator that remains valid under nonstationarity with single-trajectory data. We validate our theory with two real-world case studies. The results show that a well-calibrated TPG estimator attains low bias and variance in practical nonstationary settings.
title Estimation of Treatment Effects Under Nonstationarity via the Truncated Policy Gradient Estimator
topic Methodology
url https://arxiv.org/abs/2506.05308