Hierarchical Average-Reward Linearly-solvable Markov Decision Processes

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Infante, Guillermo, Jonsson, Anders, Gómez, Vicenç
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910519934320640
author Infante, Guillermo
Jonsson, Anders
Gómez, Vicenç
author_facet Infante, Guillermo
Jonsson, Anders
Gómez, Vicenç
contents We introduce a novel approach to hierarchical reinforcement learning for Linearly-solvable Markov Decision Processes (LMDPs) in the infinite-horizon average-reward setting. Unlike previous work, our approach allows learning low-level and high-level tasks simultaneously, without imposing limiting restrictions on the low-level tasks. Our method relies on partitions of the state space that create smaller subtasks that are easier to solve, and the equivalence between such partitions to learn more efficiently. We then exploit the compositionality of low-level tasks to exactly represent the value function of the high-level task. Experiments show that our approach can outperform flat average-reward reinforcement learning by one or several orders of magnitude.
format Preprint
id arxiv_https___arxiv_org_abs_2407_06690
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
Infante, Guillermo
Jonsson, Anders
Gómez, Vicenç
Machine Learning
Artificial Intelligence
We introduce a novel approach to hierarchical reinforcement learning for Linearly-solvable Markov Decision Processes (LMDPs) in the infinite-horizon average-reward setting. Unlike previous work, our approach allows learning low-level and high-level tasks simultaneously, without imposing limiting restrictions on the low-level tasks. Our method relies on partitions of the state space that create smaller subtasks that are easier to solve, and the equivalence between such partitions to learn more efficiently. We then exploit the compositionality of low-level tasks to exactly represent the value function of the high-level task. Experiments show that our approach can outperform flat average-reward reinforcement learning by one or several orders of magnitude.
title Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2407.06690