Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2018
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915447692066816 |
|---|---|
| author | Humayoo, Mahammad Zheng, Gengzhong Dong, Xiaoqing Miao, Liming Qiu, Shuwei Zhou, Zexun Wang, Peitao Ullah, Zakir Junejo, Naveed Ur Rehman Cheng, Xueqi |
| author_facet | Humayoo, Mahammad Zheng, Gengzhong Dong, Xiaoqing Miao, Liming Qiu, Shuwei Zhou, Zexun Wang, Peitao Ullah, Zakir Junejo, Naveed Ur Rehman Cheng, Xueqi |
| contents | Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy ($π$) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The variation between the target policy's distribution and the behavior policy's distribution can be reduced using importance sampling (IS). However, importance sampling has high variance, which is exacerbated in sequential scenarios. We propose a smooth form of importance sampling, specifically relative importance sampling (RIS), which mitigates variance and stabilizes learning. To control variance, we alter the value of the smoothness parameter $β\in[0, 1]$ in RIS. We develop the first model-free relative importance sampling off-policy actor-critic (RIS-off-PAC) algorithms in RL using this strategy. Our method uses a network to generate the target policy (actor) and evaluate the current policy ($π$) using a value function (critic) based on behavior policy samples. Our algorithms are trained using behavior policy action values in the reward function, not target policy ones. Both the actor and critic are trained using deep neural networks. Our methods performed better than or equal to several state-of-the-art RL benchmarks on OpenAI Gym challenges and synthetic datasets. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_1810_12558 |
| institution | arXiv |
| publishDate | 2018 |
| record_format | arxiv |
| spellingShingle | Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning Humayoo, Mahammad Zheng, Gengzhong Dong, Xiaoqing Miao, Liming Qiu, Shuwei Zhou, Zexun Wang, Peitao Ullah, Zakir Junejo, Naveed Ur Rehman Cheng, Xueqi Machine Learning Artificial Intelligence Off-policy learning exhibits greater instability when compared to on-policy learning in reinforcement learning (RL). The difference in probability distribution between the target policy ($π$) and the behavior policy (b) is a major cause of instability. High variance also originates from distributional mismatch. The variation between the target policy's distribution and the behavior policy's distribution can be reduced using importance sampling (IS). However, importance sampling has high variance, which is exacerbated in sequential scenarios. We propose a smooth form of importance sampling, specifically relative importance sampling (RIS), which mitigates variance and stabilizes learning. To control variance, we alter the value of the smoothness parameter $β\in[0, 1]$ in RIS. We develop the first model-free relative importance sampling off-policy actor-critic (RIS-off-PAC) algorithms in RL using this strategy. Our method uses a network to generate the target policy (actor) and evaluate the current policy ($π$) using a value function (critic) based on behavior policy samples. Our algorithms are trained using behavior policy action values in the reward function, not target policy ones. Both the actor and critic are trained using deep neural networks. Our methods performed better than or equal to several state-of-the-art RL benchmarks on OpenAI Gym challenges and synthetic datasets. |
| title | Relative Importance Sampling for off-Policy Actor-Critic in Deep Reinforcement Learning |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/1810.12558 |