Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911627294539776 |
|---|---|
| author | Benac, Leo Sharma, Abhishek Parbhoo, Sonali Doshi-Velez, Finale |
| author_facet | Benac, Leo Sharma, Abhishek Parbhoo, Sonali Doshi-Velez, Finale |
| contents | We consider the problem of estimating the transition dynamics $T^*$ from near-optimal expert trajectories in the context of offline model-based reinforcement learning. We develop a novel constraint-based method, Inverse Transition Learning, that treats the limited coverage of the expert trajectories as a \emph{feature}: we use the fact that the expert is near-optimal to inform our estimate of $T^*$. We integrate our constraints into a Bayesian approach. Across both synthetic environments and real healthcare scenarios like Intensive Care Unit (ICU) patient management in hypotension, we demonstrate not only significant improvements in decision-making, but that our posterior can inform when transfer will be successful. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_05174 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories Benac, Leo Sharma, Abhishek Parbhoo, Sonali Doshi-Velez, Finale Machine Learning Artificial Intelligence We consider the problem of estimating the transition dynamics $T^*$ from near-optimal expert trajectories in the context of offline model-based reinforcement learning. We develop a novel constraint-based method, Inverse Transition Learning, that treats the limited coverage of the expert trajectories as a \emph{feature}: we use the fact that the expert is near-optimal to inform our estimate of $T^*$. We integrate our constraints into a Bayesian approach. Across both synthetic environments and real healthcare scenarios like Intensive Care Unit (ICU) patient management in hypotension, we demonstrate not only significant improvements in decision-making, but that our posterior can inform when transfer will be successful. |
| title | Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2411.05174 |