Efficient Inference after Directionally Stable Adaptive Experiments

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Zikai, Zenati, Houssam, Kallus, Nathan, Gretton, Arthur, Khamaru, Koulik, Bibaut, Aurélien
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911467317493760
author Shen, Zikai
Zenati, Houssam
Kallus, Nathan
Gretton, Arthur
Khamaru, Koulik
Bibaut, Aurélien
author_facet Shen, Zikai
Zenati, Houssam
Kallus, Nathan
Gretton, Arthur
Khamaru, Koulik
Bibaut, Aurélien
contents We study inference on scalar-valued pathwise differentiable targets after adaptive data collection, such as a bandit algorithm. We introduce a novel target-specific condition, directional stability, which is strictly weaker than previously imposed target-agnostic stability conditions. Under directional stability, we show that estimators that would have been efficient under i.i.d. data remain asymptotically normal and semiparametrically efficient when computed from adaptively collected trajectories. The canonical gradient has a martingale form, and directional stability guarantees stabilization of its predictable quadratic variation, enabling high-dimensional asymptotic normality. We characterize efficiency using a convolution theorem for the adaptive-data setting, and give a condition under which the one-step estimator attains the efficiency bound. We verify directional stability for LinUCB, yielding the first semiparametric efficiency guarantee for a regular scalar target under LinUCB sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21478
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficient Inference after Directionally Stable Adaptive Experiments
Shen, Zikai
Zenati, Houssam
Kallus, Nathan
Gretton, Arthur
Khamaru, Koulik
Bibaut, Aurélien
Machine Learning
Statistics Theory
Methodology
We study inference on scalar-valued pathwise differentiable targets after adaptive data collection, such as a bandit algorithm. We introduce a novel target-specific condition, directional stability, which is strictly weaker than previously imposed target-agnostic stability conditions. Under directional stability, we show that estimators that would have been efficient under i.i.d. data remain asymptotically normal and semiparametrically efficient when computed from adaptively collected trajectories. The canonical gradient has a martingale form, and directional stability guarantees stabilization of its predictable quadratic variation, enabling high-dimensional asymptotic normality. We characterize efficiency using a convolution theorem for the adaptive-data setting, and give a condition under which the one-step estimator attains the efficiency bound. We verify directional stability for LinUCB, yielding the first semiparametric efficiency guarantee for a regular scalar target under LinUCB sampling.
title Efficient Inference after Directionally Stable Adaptive Experiments
topic Machine Learning
Statistics Theory
Methodology
url https://arxiv.org/abs/2602.21478