Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Yuhan, Han, Eugene, Hu, Yifan, Zhou, Wenzhuo, Qi, Zhengling, Cui, Yifan, Zhu, Ruoqing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918005760327680
author Li, Yuhan
Han, Eugene
Hu, Yifan
Zhou, Wenzhuo
Qi, Zhengling
Cui, Yifan
Zhu, Ruoqing
author_facet Li, Yuhan
Han, Eugene
Hu, Yifan
Zhou, Wenzhuo
Qi, Zhengling
Cui, Yifan
Zhu, Ruoqing
contents This paper addresses the challenge of offline policy learning in reinforcement learning with continuous action spaces when unmeasured confounders are present. While most existing research focuses on policy evaluation within partially observable Markov decision processes (POMDPs) and assumes discrete action spaces, we advance this field by establishing a novel identification result to enable the nonparametric estimation of policy value for a given target policy under an infinite-horizon framework. Leveraging this identification, we develop a minimax estimator and introduce a policy-gradient-based algorithm to identify the in-class optimal policy that maximizes the estimated policy value. Furthermore, we provide theoretical results regarding the consistency, finite-sample error bound, and regret bound of the resulting optimal policy. Extensive simulations and a real-world application using the German Family Panel data demonstrate the effectiveness of our proposed methodology.
format Preprint
id arxiv_https___arxiv_org_abs_2505_00304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
Li, Yuhan
Han, Eugene
Hu, Yifan
Zhou, Wenzhuo
Qi, Zhengling
Cui, Yifan
Zhu, Ruoqing
Machine Learning
Methodology
This paper addresses the challenge of offline policy learning in reinforcement learning with continuous action spaces when unmeasured confounders are present. While most existing research focuses on policy evaluation within partially observable Markov decision processes (POMDPs) and assumes discrete action spaces, we advance this field by establishing a novel identification result to enable the nonparametric estimation of policy value for a given target policy under an infinite-horizon framework. Leveraging this identification, we develop a minimax estimator and introduce a policy-gradient-based algorithm to identify the in-class optimal policy that maximizes the estimated policy value. Furthermore, we provide theoretical results regarding the consistency, finite-sample error bound, and regret bound of the resulting optimal policy. Extensive simulations and a real-world application using the German Family Panel data demonstrate the effectiveness of our proposed methodology.
title Reinforcement Learning with Continuous Actions Under Unmeasured Confounding
topic Machine Learning
Methodology
url https://arxiv.org/abs/2505.00304