Model predictive control-based value estimation for efficient reinforcement learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Qizhen, Liu, Kexin, Chen, Lei
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916452558176256
author Wu, Qizhen
Liu, Kexin
Chen, Lei
author_facet Wu, Qizhen
Liu, Kexin
Chen, Lei
contents Reinforcement learning suffers from limitations in real practices primarily due to the number of required interactions with virtual environments. It results in a challenging problem because we are implausible to obtain a local optimal strategy with only a few attempts for many learning methods. Hereby, we design an improved reinforcement learning method based on model predictive control that models the environment through a data-driven approach. Based on the learned environment model, it performs multi-step prediction to estimate the value function and optimize the policy. The method demonstrates higher learning efficiency, faster convergent speed of strategies tending to the local optimal value, and less sample capacity space required by experience replay buffers. Experimental results, both in classic databases and in a dynamic obstacle avoidance scenario for an unmanned aerial vehicle, validate the proposed approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16646
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Model predictive control-based value estimation for efficient reinforcement learning
Wu, Qizhen
Liu, Kexin
Chen, Lei
Machine Learning
Reinforcement learning suffers from limitations in real practices primarily due to the number of required interactions with virtual environments. It results in a challenging problem because we are implausible to obtain a local optimal strategy with only a few attempts for many learning methods. Hereby, we design an improved reinforcement learning method based on model predictive control that models the environment through a data-driven approach. Based on the learned environment model, it performs multi-step prediction to estimate the value function and optimize the policy. The method demonstrates higher learning efficiency, faster convergent speed of strategies tending to the local optimal value, and less sample capacity space required by experience replay buffers. Experimental results, both in classic databases and in a dynamic obstacle avoidance scenario for an unmanned aerial vehicle, validate the proposed approaches.
title Model predictive control-based value estimation for efficient reinforcement learning
topic Machine Learning
url https://arxiv.org/abs/2310.16646