Informed POMDP: Leveraging Additional Information in Model-Based RL
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913879312826368 |
|---|---|
| author | Lambrechts, Gaspard Bolland, Adrien Ernst, Damien |
| author_facet | Lambrechts, Gaspard Bolland, Adrien Ernst, Damien |
| contents | In this work, we generalize the problem of learning through interaction in a POMDP by accounting for eventual additional information available at training time. First, we introduce the informed POMDP, a new learning paradigm offering a clear distinction between the information at training and the observation at execution. Next, we propose an objective that leverages this information for learning a sufficient statistic of the history for the optimal control. We then adapt this informed objective to learn a world model able to sample latent trajectories. Finally, we empirically show a learning speed improvement in several environments using this informed world model in the Dreamer algorithm. These results and the simplicity of the proposed adaptation advocate for a systematic consideration of eventual additional information when learning in a POMDP using model-based RL. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_11488 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Informed POMDP: Leveraging Additional Information in Model-Based RL Lambrechts, Gaspard Bolland, Adrien Ernst, Damien Machine Learning In this work, we generalize the problem of learning through interaction in a POMDP by accounting for eventual additional information available at training time. First, we introduce the informed POMDP, a new learning paradigm offering a clear distinction between the information at training and the observation at execution. Next, we propose an objective that leverages this information for learning a sufficient statistic of the history for the optimal control. We then adapt this informed objective to learn a world model able to sample latent trajectories. Finally, we empirically show a learning speed improvement in several environments using this informed world model in the Dreamer algorithm. These results and the simplicity of the proposed adaptation advocate for a systematic consideration of eventual additional information when learning in a POMDP using model-based RL. |
| title | Informed POMDP: Leveraging Additional Information in Model-Based RL |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2306.11488 |