Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lyu, Xubo, Li, Site, Siriya, Seth, Pu, Ye, Chen, Mo
Format: Preprint
Published: 2020
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929577751740416
author Lyu, Xubo
Li, Site
Siriya, Seth
Pu, Ye
Chen, Mo
author_facet Lyu, Xubo
Li, Site
Siriya, Seth
Pu, Ye
Chen, Mo
contents In this paper, a novel optimal control-based baseline function is presented for the policy gradient method in deep reinforcement learning (RL). The baseline is obtained by computing the value function of an optimal control problem, which is formed to be closely associated with the RL task. In contrast to the traditional baseline aimed at variance reduction of policy gradient estimates, our work utilizes the optimal control value function to introduce a novel aspect to the role of baseline -- providing guided exploration during policy learning. This aspect is less discussed in prior works. We validate our baseline on robot learning tasks, showing its effectiveness in guided exploration, particularly in sparse reward environments.
format Preprint
id arxiv_https___arxiv_org_abs_2011_02073
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
Lyu, Xubo
Li, Site
Siriya, Seth
Pu, Ye
Chen, Mo
Machine Learning
Artificial Intelligence
Robotics
Systems and Control
In this paper, a novel optimal control-based baseline function is presented for the policy gradient method in deep reinforcement learning (RL). The baseline is obtained by computing the value function of an optimal control problem, which is formed to be closely associated with the RL task. In contrast to the traditional baseline aimed at variance reduction of policy gradient estimates, our work utilizes the optimal control value function to introduce a novel aspect to the role of baseline -- providing guided exploration during policy learning. This aspect is less discussed in prior works. We validate our baseline on robot learning tasks, showing its effectiveness in guided exploration, particularly in sparse reward environments.
title Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
topic Machine Learning
Artificial Intelligence
Robotics
Systems and Control
url https://arxiv.org/abs/2011.02073