Model-based Reinforcement Learning for Parameterized Action Spaces

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Renhao, Fu, Haotian, Miao, Yilin, Konidaris, George
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911885493796864
author Zhang, Renhao
Fu, Haotian
Miao, Yilin
Konidaris, George
author_facet Zhang, Renhao
Fu, Haotian
Miao, Yilin
Konidaris, George
contents We propose a novel model-based reinforcement learning algorithm -- Dynamics Learning and predictive control with Parameterized Actions (DLPA) -- for Parameterized Action Markov Decision Processes (PAMDPs). The agent learns a parameterized-action-conditioned dynamics model and plans with a modified Model Predictive Path Integral control. We theoretically quantify the difference between the generated trajectory and the optimal trajectory during planning in terms of the value they achieved through the lens of Lipschitz Continuity. Our empirical results on several standard benchmarks show that our algorithm achieves superior sample efficiency and asymptotic performance than state-of-the-art PAMDP methods.
format Preprint
id arxiv_https___arxiv_org_abs_2404_03037
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Model-based Reinforcement Learning for Parameterized Action Spaces
Zhang, Renhao
Fu, Haotian
Miao, Yilin
Konidaris, George
Machine Learning
Artificial Intelligence
We propose a novel model-based reinforcement learning algorithm -- Dynamics Learning and predictive control with Parameterized Actions (DLPA) -- for Parameterized Action Markov Decision Processes (PAMDPs). The agent learns a parameterized-action-conditioned dynamics model and plans with a modified Model Predictive Path Integral control. We theoretically quantify the difference between the generated trajectory and the optimal trajectory during planning in terms of the value they achieved through the lens of Lipschitz Continuity. Our empirical results on several standard benchmarks show that our algorithm achieves superior sample efficiency and asymptotic performance than state-of-the-art PAMDP methods.
title Model-based Reinforcement Learning for Parameterized Action Spaces
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2404.03037