Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Bai, Chenjia, Liu, Peng, Liu, Kaiyu, Wang, Lingxiao, Zhao, Yingnan, Han, Lei
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909156269621248
author Bai, Chenjia
Liu, Peng
Liu, Kaiyu
Wang, Lingxiao
Zhao, Yingnan
Han, Lei
author_facet Bai, Chenjia
Liu, Peng
Liu, Kaiyu
Wang, Lingxiao
Zhao, Yingnan
Han, Lei
contents Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show promising results in simple environments but often get stuck in environments with multimodal and stochastic dynamics. In this work, we propose a variational dynamic model based on the conditional variational inference to model the multimodality and stochasticity. We consider the environmental state-action transition as a conditional generative process by generating the next-state prediction under the condition of the current state, action, and latent variable, which provides a better understanding of the dynamics and leads a better performance in exploration. We derive an upper bound of the negative log-likelihood of the environmental transition and use such an upper bound as the intrinsic reward for exploration, which allows the agent to learn skills by self-supervised exploration without observing extrinsic rewards. We evaluate the proposed method on several image-based simulation tasks and a real robotic manipulating task. Our method outperforms several state-of-the-art environment model-based exploration approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2010_08755
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
Bai, Chenjia
Liu, Peng
Liu, Kaiyu
Wang, Lingxiao
Zhao, Yingnan
Han, Lei
Machine Learning
Computer Vision and Pattern Recognition
Robotics
Efficient exploration remains a challenging problem in reinforcement learning, especially for tasks where extrinsic rewards from environments are sparse or even totally disregarded. Significant advances based on intrinsic motivation show promising results in simple environments but often get stuck in environments with multimodal and stochastic dynamics. In this work, we propose a variational dynamic model based on the conditional variational inference to model the multimodality and stochasticity. We consider the environmental state-action transition as a conditional generative process by generating the next-state prediction under the condition of the current state, action, and latent variable, which provides a better understanding of the dynamics and leads a better performance in exploration. We derive an upper bound of the negative log-likelihood of the environmental transition and use such an upper bound as the intrinsic reward for exploration, which allows the agent to learn skills by self-supervised exploration without observing extrinsic rewards. We evaluate the proposed method on several image-based simulation tasks and a real robotic manipulating task. Our method outperforms several state-of-the-art environment model-based exploration approaches.
title Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
topic Machine Learning
Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2010.08755