POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Ruijia, Zhang, Xiangyu, Qi, Zhengling, Wu, Yue, Xu, Yanxun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914288208183296
author Zhang, Ruijia
Zhang, Xiangyu
Qi, Zhengling
Wu, Yue
Xu, Yanxun
author_facet Zhang, Ruijia
Zhang, Xiangyu
Qi, Zhengling
Wu, Yue
Xu, Yanxun
contents Dynamic treatment regimes (DTRs) provide a principled framework for optimizing sequential decision-making in domains where decisions must adapt over time in response to individual trajectories, such as healthcare, education, and digital interventions. However, existing statistical methods often rely on strong positivity assumptions and lack robustness under partial data coverage, while offline reinforcement learning approaches typically focus on average training performance, lack statistical guarantees, and require solving complex optimization problems. To address these challenges, we propose POLAR, a novel pessimistic model-based policy learning algorithm for offline DTR optimization. POLAR estimates the transition dynamics from offline data and quantifies uncertainty for each history-action pair. A pessimistic penalty is then incorporated into the reward function to discourage actions with high uncertainty. Unlike many existing methods that focus on average training performance or provide guarantees only for an oracle policy, POLAR directly targets the suboptimality of the final learned policy and offers theoretical guarantees, without relying on computationally intensive minimax or constrained optimization procedures. To the best of our knowledge, POLAR is the first model-based DTR method to provide both statistical and computational guarantees, including finite-sample bounds on policy suboptimality. Empirical results on both synthetic data and the MIMIC-III dataset demonstrate that POLAR outperforms state-of-the-art methods and yields near-optimal, history-aware treatment strategies.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20406
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
Zhang, Ruijia
Zhang, Xiangyu
Qi, Zhengling
Wu, Yue
Xu, Yanxun
Machine Learning
Information Theory
Methodology
Dynamic treatment regimes (DTRs) provide a principled framework for optimizing sequential decision-making in domains where decisions must adapt over time in response to individual trajectories, such as healthcare, education, and digital interventions. However, existing statistical methods often rely on strong positivity assumptions and lack robustness under partial data coverage, while offline reinforcement learning approaches typically focus on average training performance, lack statistical guarantees, and require solving complex optimization problems. To address these challenges, we propose POLAR, a novel pessimistic model-based policy learning algorithm for offline DTR optimization. POLAR estimates the transition dynamics from offline data and quantifies uncertainty for each history-action pair. A pessimistic penalty is then incorporated into the reward function to discourage actions with high uncertainty. Unlike many existing methods that focus on average training performance or provide guarantees only for an oracle policy, POLAR directly targets the suboptimality of the final learned policy and offers theoretical guarantees, without relying on computationally intensive minimax or constrained optimization procedures. To the best of our knowledge, POLAR is the first model-based DTR method to provide both statistical and computational guarantees, including finite-sample bounds on policy suboptimality. Empirical results on both synthetic data and the MIMIC-III dataset demonstrate that POLAR outperforms state-of-the-art methods and yields near-optimal, history-aware treatment strategies.
title POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
topic Machine Learning
Information Theory
Methodology
url https://arxiv.org/abs/2506.20406