Central Limit Theorems for Transition Probabilities of Controlled Markov Chains

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Su, Ziwei, Banerjee, Imon, Klabjan, Diego
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908911115698176
author Su, Ziwei
Banerjee, Imon
Klabjan, Diego
author_facet Su, Ziwei
Banerjee, Imon
Klabjan, Diego
contents We develop a central limit theorem (CLT) for a non-parametric estimator of the transition matrices in controlled Markov chains (CMCs) with finite state-action spaces. Our results establish precise conditions on the logging policy under which the estimator is asymptotically normal, and reveal settings in which no CLT can exist. We then build on it to derive CLTs for the value, Q-, and advantage functions of any stationary stochastic policy, including the optimal policy recovered from the estimated model. Goodness-of-fit tests are derived as a corollary, which enable to test whether the logged data is stochastic. These results provide new statistical tools for offline policy evaluation and optimal policy recovery, and enable hypothesis tests for transition probabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01517
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Central Limit Theorems for Transition Probabilities of Controlled Markov Chains
Su, Ziwei
Banerjee, Imon
Klabjan, Diego
Statistics Theory
Probability
Machine Learning
Primary 60F05, Secondary 60J05, 62M05, 93E20
G.3; I.2.6; I.2.8
We develop a central limit theorem (CLT) for a non-parametric estimator of the transition matrices in controlled Markov chains (CMCs) with finite state-action spaces. Our results establish precise conditions on the logging policy under which the estimator is asymptotically normal, and reveal settings in which no CLT can exist. We then build on it to derive CLTs for the value, Q-, and advantage functions of any stationary stochastic policy, including the optimal policy recovered from the estimated model. Goodness-of-fit tests are derived as a corollary, which enable to test whether the logged data is stochastic. These results provide new statistical tools for offline policy evaluation and optimal policy recovery, and enable hypothesis tests for transition probabilities.
title Central Limit Theorems for Transition Probabilities of Controlled Markov Chains
topic Statistics Theory
Probability
Machine Learning
Primary 60F05, Secondary 60J05, 62M05, 93E20
G.3; I.2.6; I.2.8
url https://arxiv.org/abs/2508.01517