Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guan, Yang, Duan, Jingliang, Li, Shengbo Eben, Li, Jie, Chen, Jianyu, Cheng, Bo
Natura: Preprint
Pubblicazione: 2021
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911783147536384
author Guan, Yang
Duan, Jingliang
Li, Shengbo Eben
Li, Jie
Chen, Jianyu
Cheng, Bo
author_facet Guan, Yang
Duan, Jingliang
Li, Shengbo Eben
Li, Jie
Chen, Jianyu
Cheng, Bo
contents Reinforcement learning (RL) shows great potential in sequential decision-making. At present, mainstream RL algorithms are data-driven, which usually yield better asymptotic performance but much slower convergence compared with model-driven methods. This paper proposes mixed policy gradient (MPG) algorithm, which fuses the empirical data and the transition model in policy gradient (PG) to accelerate convergence without performance degradation. Formally, MPG is constructed as a weighted average of the data-driven and model-driven PGs, where the former is the derivative of the learned Q-value function, and the latter is that of the model-predictive return. To guide the weight design, we analyze and compare the upper bound of each PG error. Relying on that, a rule-based method is employed to heuristically adjust the weights. In particular, to get a better PG, the weight of the data-driven PG is designed to grow along the learning process while the other to decrease. Simulation results show that the MPG method achieves the best asymptotic performance and convergence speed compared with other baseline algorithms.
format Preprint
id arxiv_https___arxiv_org_abs_2102_11513
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
Guan, Yang
Duan, Jingliang
Li, Shengbo Eben
Li, Jie
Chen, Jianyu
Cheng, Bo
Machine Learning
Reinforcement learning (RL) shows great potential in sequential decision-making. At present, mainstream RL algorithms are data-driven, which usually yield better asymptotic performance but much slower convergence compared with model-driven methods. This paper proposes mixed policy gradient (MPG) algorithm, which fuses the empirical data and the transition model in policy gradient (PG) to accelerate convergence without performance degradation. Formally, MPG is constructed as a weighted average of the data-driven and model-driven PGs, where the former is the derivative of the learned Q-value function, and the latter is that of the model-predictive return. To guide the weight design, we analyze and compare the upper bound of each PG error. Relying on that, a rule-based method is employed to heuristically adjust the weights. In particular, to get a better PG, the weight of the data-driven PG is designed to grow along the learning process while the other to decrease. Simulation results show that the MPG method achieves the best asymptotic performance and convergence speed compared with other baseline algorithms.
title Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
topic Machine Learning
url https://arxiv.org/abs/2102.11513