A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Palash, Wang, Xinru, Nalamada, Trikay, Agarwal, Shruti, Jahja, Maria, Chakraborty, Bibhas
Format: Preprint
Published: 2021
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909414268600320
author Ghosh, Palash
Wang, Xinru
Nalamada, Trikay
Agarwal, Shruti
Jahja, Maria
Chakraborty, Bibhas
author_facet Ghosh, Palash
Wang, Xinru
Nalamada, Trikay
Agarwal, Shruti
Jahja, Maria
Chakraborty, Bibhas
contents A dynamic treatment regimen (DTR) is a set of decision rules to personalize treatments for an individual using their medical history. The Q-learning-based Q-shared algorithm has been used to develop DTRs that involve decision rules shared across multiple stages of intervention. We show that the existing Q-shared algorithm can suffer from non-convergence due to the use of linear models in the Q-learning setup, and identify the condition under which Q-shared fails. We develop a penalized Q-shared algorithm that not only converges in settings that violate the condition, but can outperform the original Q-shared algorithm even when the condition is satisfied. We give evidence for the proposed method in a real-world application and several synthetic simulations.
format Preprint
id arxiv_https___arxiv_org_abs_2107_07875
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens
Ghosh, Palash
Wang, Xinru
Nalamada, Trikay
Agarwal, Shruti
Jahja, Maria
Chakraborty, Bibhas
Machine Learning
A dynamic treatment regimen (DTR) is a set of decision rules to personalize treatments for an individual using their medical history. The Q-learning-based Q-shared algorithm has been used to develop DTRs that involve decision rules shared across multiple stages of intervention. We show that the existing Q-shared algorithm can suffer from non-convergence due to the use of linear models in the Q-learning setup, and identify the condition under which Q-shared fails. We develop a penalized Q-shared algorithm that not only converges in settings that violate the condition, but can outperform the original Q-shared algorithm even when the condition is satisfied. We give evidence for the proposed method in a real-world application and several synthetic simulations.
title A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens
topic Machine Learning
url https://arxiv.org/abs/2107.07875