Stochastic Primal-Dual Q-Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeong, Narim, Lee, Donghwan, He, Niao
Format: Preprint
Published: 2018
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909693584080896
author Jeong, Narim
Lee, Donghwan
He, Niao
author_facet Jeong, Narim
Lee, Donghwan
He, Niao
contents In this work, we present a new model-free and off-policy reinforcement learning (RL) algorithm, that is capable of finding a near-optimal policy with state-action observations from arbitrary behavior policies. Our algorithm, called the stochastic primal-dual Q-learning (SPD Q-learning), hinges upon a new linear programming formulation and a dual perspective of the standard Q-learning. In contrast to previous primal-dual RL algorithms, the SPD Q-learning includes a Q-function estimation step, thus allowing to recover an approximate policy from the primal solution as well as the dual solution. We prove a first-of-its-kind result that the SPD Q-learning guarantees a certain convergence rate, even when the state-action distribution is time-varying but sub-linearly converges to a stationary distribution. Numerical experiments are provided to demonstrate the off-policy learning abilities of the proposed algorithm in comparison to the standard Q-learning.
format Preprint
id arxiv_https___arxiv_org_abs_1810_08298
institution arXiv
publishDate 2018
record_format arxiv
spellingShingle Stochastic Primal-Dual Q-Learning
Jeong, Narim
Lee, Donghwan
He, Niao
Optimization and Control
In this work, we present a new model-free and off-policy reinforcement learning (RL) algorithm, that is capable of finding a near-optimal policy with state-action observations from arbitrary behavior policies. Our algorithm, called the stochastic primal-dual Q-learning (SPD Q-learning), hinges upon a new linear programming formulation and a dual perspective of the standard Q-learning. In contrast to previous primal-dual RL algorithms, the SPD Q-learning includes a Q-function estimation step, thus allowing to recover an approximate policy from the primal solution as well as the dual solution. We prove a first-of-its-kind result that the SPD Q-learning guarantees a certain convergence rate, even when the state-action distribution is time-varying but sub-linearly converges to a stationary distribution. Numerical experiments are provided to demonstrate the off-policy learning abilities of the proposed algorithm in comparison to the standard Q-learning.
title Stochastic Primal-Dual Q-Learning
topic Optimization and Control
url https://arxiv.org/abs/1810.08298