$λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Voelcker, Claas A, Ahmadian, Arash, Abachi, Romina, Gilitschenski, Igor, Farahmand, Amir-massoud
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929259731222528
author Voelcker, Claas A
Ahmadian, Arash
Abachi, Romina
Gilitschenski, Igor
Farahmand, Amir-massoud
author_facet Voelcker, Claas A
Ahmadian, Arash
Abachi, Romina
Gilitschenski, Igor
Farahmand, Amir-massoud
contents The idea of decision-aware model learning, that models should be accurate where it matters for decision-making, has gained prominence in model-based reinforcement learning. While promising theoretical results have been established, the empirical performance of algorithms leveraging a decision-aware loss has been lacking, especially in continuous control problems. In this paper, we present a study on the necessary components for decision-aware reinforcement learning models and we showcase design choices that enable well-performing algorithms. To this end, we provide a theoretical and empirical investigation into algorithmic ideas in the field. We highlight that empirical design decisions established in the MuZero line of works, most importantly the use of a latent model, are vital to achieving good performance for related algorithms. Furthermore, we show that the MuZero loss function is biased in stochastic environments and establish that this bias has practical consequences. Building on these findings, we present an overview of which decision-aware loss functions are best used in what empirical scenarios, providing actionable insights to practitioners in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2306_17366
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle $λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
Voelcker, Claas A
Ahmadian, Arash
Abachi, Romina
Gilitschenski, Igor
Farahmand, Amir-massoud
Machine Learning
Artificial Intelligence
The idea of decision-aware model learning, that models should be accurate where it matters for decision-making, has gained prominence in model-based reinforcement learning. While promising theoretical results have been established, the empirical performance of algorithms leveraging a decision-aware loss has been lacking, especially in continuous control problems. In this paper, we present a study on the necessary components for decision-aware reinforcement learning models and we showcase design choices that enable well-performing algorithms. To this end, we provide a theoretical and empirical investigation into algorithmic ideas in the field. We highlight that empirical design decisions established in the MuZero line of works, most importantly the use of a latent model, are vital to achieving good performance for related algorithms. Furthermore, we show that the MuZero loss function is biased in stochastic environments and establish that this bias has practical consequences. Building on these findings, we present an overview of which decision-aware loss functions are best used in what empirical scenarios, providing actionable insights to practitioners in the field.
title $λ$-models: Effective Decision-Aware Reinforcement Learning with Latent Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2306.17366