Informed POMDP: Leveraging Additional Information in Model-Based RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lambrechts, Gaspard, Bolland, Adrien, Ernst, Damien
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913879312826368
author Lambrechts, Gaspard
Bolland, Adrien
Ernst, Damien
author_facet Lambrechts, Gaspard
Bolland, Adrien
Ernst, Damien
contents In this work, we generalize the problem of learning through interaction in a POMDP by accounting for eventual additional information available at training time. First, we introduce the informed POMDP, a new learning paradigm offering a clear distinction between the information at training and the observation at execution. Next, we propose an objective that leverages this information for learning a sufficient statistic of the history for the optimal control. We then adapt this informed objective to learn a world model able to sample latent trajectories. Finally, we empirically show a learning speed improvement in several environments using this informed world model in the Dreamer algorithm. These results and the simplicity of the proposed adaptation advocate for a systematic consideration of eventual additional information when learning in a POMDP using model-based RL.
format Preprint
id arxiv_https___arxiv_org_abs_2306_11488
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Informed POMDP: Leveraging Additional Information in Model-Based RL
Lambrechts, Gaspard
Bolland, Adrien
Ernst, Damien
Machine Learning
In this work, we generalize the problem of learning through interaction in a POMDP by accounting for eventual additional information available at training time. First, we introduce the informed POMDP, a new learning paradigm offering a clear distinction between the information at training and the observation at execution. Next, we propose an objective that leverages this information for learning a sufficient statistic of the history for the optimal control. We then adapt this informed objective to learn a world model able to sample latent trajectories. Finally, we empirically show a learning speed improvement in several environments using this informed world model in the Dreamer algorithm. These results and the simplicity of the proposed adaptation advocate for a systematic consideration of eventual additional information when learning in a POMDP using model-based RL.
title Informed POMDP: Leveraging Additional Information in Model-Based RL
topic Machine Learning
url https://arxiv.org/abs/2306.11488