A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909414268600320 |
|---|---|
| author | Ghosh, Palash Wang, Xinru Nalamada, Trikay Agarwal, Shruti Jahja, Maria Chakraborty, Bibhas |
| author_facet | Ghosh, Palash Wang, Xinru Nalamada, Trikay Agarwal, Shruti Jahja, Maria Chakraborty, Bibhas |
| contents | A dynamic treatment regimen (DTR) is a set of decision rules to personalize treatments for an individual using their medical history. The Q-learning-based Q-shared algorithm has been used to develop DTRs that involve decision rules shared across multiple stages of intervention. We show that the existing Q-shared algorithm can suffer from non-convergence due to the use of linear models in the Q-learning setup, and identify the condition under which Q-shared fails. We develop a penalized Q-shared algorithm that not only converges in settings that violate the condition, but can outperform the original Q-shared algorithm even when the condition is satisfied. We give evidence for the proposed method in a real-world application and several synthetic simulations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2107_07875 |
| institution | arXiv |
| publishDate | 2021 |
| record_format | arxiv |
| spellingShingle | A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens Ghosh, Palash Wang, Xinru Nalamada, Trikay Agarwal, Shruti Jahja, Maria Chakraborty, Bibhas Machine Learning A dynamic treatment regimen (DTR) is a set of decision rules to personalize treatments for an individual using their medical history. The Q-learning-based Q-shared algorithm has been used to develop DTRs that involve decision rules shared across multiple stages of intervention. We show that the existing Q-shared algorithm can suffer from non-convergence due to the use of linear models in the Q-learning setup, and identify the condition under which Q-shared fails. We develop a penalized Q-shared algorithm that not only converges in settings that violate the condition, but can outperform the original Q-shared algorithm even when the condition is satisfied. We give evidence for the proposed method in a real-world application and several synthetic simulations. |
| title | A Penalized Shared-parameter Algorithm for Estimating Optimal Dynamic Treatment Regimens |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2107.07875 |