Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Humayoo, Mahammad, Zheng, Gengzhong, Dong, Xiaoqing, Miao, Liming, Qiu, Shuwei, Zhou, Zexun, Wang, Peitao, Ullah, Zakir, Junejo, Naveed Ur Rehman, Cheng, Xueqi
Format: Preprint
Published: 2018
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915447692066816
author Humayoo, Mahammad
Zheng, Gengzhong
Dong, Xiaoqing
Miao, Liming
Qiu, Shuwei
Zhou, Zexun
Wang, Peitao
Ullah, Zakir
Junejo, Naveed Ur Rehman
Cheng, Xueqi
author_facet Humayoo, Mahammad
Zheng, Gengzhong
Dong, Xiaoqing
Miao, Liming
Qiu, Shuwei
Zhou, Zexun
Wang, Peitao
Ullah, Zakir
Junejo, Naveed Ur Rehman
Cheng, Xueqi
contents Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy ($π$) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The variation between the target policy's distribution and the behavior policy's distribution can be reduced using importance sampling (IS). However, importance sampling has high variance, which is exacerbated in sequential scenarios. We propose a smooth form of importance sampling, specifically relative importance sampling (RIS), which mitigates variance and stabilizes learning. To control variance, we alter the value of the smoothness parameter $β\in[0, 1]$ in RIS. We develop the first model-free relative importance sampling off-policy actor-critic (RIS-off-PAC) algorithms in RL using this strategy. Our method uses a network to generate the target policy (actor) and evaluate the current policy ($π$) using a value function (critic) based on behavior policy samples. Our algorithms are trained using behavior policy action values in the reward function, not target policy ones. Both the actor and critic are trained using deep neural networks. Our methods performed better than or equal to several state-of-the-art RL benchmarks on OpenAI Gym challenges and synthetic datasets.
format Preprint
id arxiv_https___arxiv_org_abs_1810_12558
institution arXiv
publishDate 2018
record_format arxiv
spellingShingle Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
Humayoo, Mahammad
Zheng, Gengzhong
Dong, Xiaoqing
Miao, Liming
Qiu, Shuwei
Zhou, Zexun
Wang, Peitao
Ullah, Zakir
Junejo, Naveed Ur Rehman
Cheng, Xueqi
Machine Learning
Artificial Intelligence
Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy ($π$) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The variation between the target policy's distribution and the behavior policy's distribution can be reduced using importance sampling (IS). However, importance sampling has high variance, which is exacerbated in sequential scenarios. We propose a smooth form of importance sampling, specifically relative importance sampling (RIS), which mitigates variance and stabilizes learning. To control variance, we alter the value of the smoothness parameter $β\in[0, 1]$ in RIS. We develop the first model-free relative importance sampling off-policy actor-critic (RIS-off-PAC) algorithms in RL using this strategy. Our method uses a network to generate the target policy (actor) and evaluate the current policy ($π$) using a value function (critic) based on behavior policy samples. Our algorithms are trained using behavior policy action values in the reward function, not target policy ones. Both the actor and critic are trained using deep neural networks. Our methods performed better than or equal to several state-of-the-art RL benchmarks on OpenAI Gym challenges and synthetic datasets.
title Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/1810.12558