Upside-Down Reinforcement Learning for More Interpretable Optimal Control

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cardenas-Cartagena, Juan, Falzari, Massimiliano, Zullich, Marco, Sabatelli, Matthia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929595395080192
author Cardenas-Cartagena, Juan
Falzari, Massimiliano
Zullich, Marco
Sabatelli, Matthia
author_facet Cardenas-Cartagena, Juan
Falzari, Massimiliano
Zullich, Marco
Sabatelli, Matthia
contents Model-Free Reinforcement Learning (RL) algorithms either learn how to map states to expected rewards or search for policies that can maximize a certain performance function. Model-Based algorithms instead, aim to learn an approximation of the underlying model of the RL environment and then use it in combination with planning algorithms. Upside-Down Reinforcement Learning (UDRL) is a novel learning paradigm that aims to learn how to predict actions from states and desired commands. This task is formulated as a Supervised Learning problem and has successfully been tackled by Neural Networks (NNs). In this paper, we investigate whether function approximation algorithms other than NNs can also be used within a UDRL framework. Our experiments, performed over several popular optimal control benchmarks, show that tree-based methods like Random Forests and Extremely Randomized Trees can perform just as well as NNs with the significant benefit of resulting in policies that are inherently more interpretable than NNs, therefore paving the way for more transparent, safe, and robust RL.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11457
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Upside-Down Reinforcement Learning for More Interpretable Optimal Control
Cardenas-Cartagena, Juan
Falzari, Massimiliano
Zullich, Marco
Sabatelli, Matthia
Machine Learning
Model-Free Reinforcement Learning (RL) algorithms either learn how to map states to expected rewards or search for policies that can maximize a certain performance function. Model-Based algorithms instead, aim to learn an approximation of the underlying model of the RL environment and then use it in combination with planning algorithms. Upside-Down Reinforcement Learning (UDRL) is a novel learning paradigm that aims to learn how to predict actions from states and desired commands. This task is formulated as a Supervised Learning problem and has successfully been tackled by Neural Networks (NNs). In this paper, we investigate whether function approximation algorithms other than NNs can also be used within a UDRL framework. Our experiments, performed over several popular optimal control benchmarks, show that tree-based methods like Random Forests and Extremely Randomized Trees can perform just as well as NNs with the significant benefit of resulting in policies that are inherently more interpretable than NNs, therefore paving the way for more transparent, safe, and robust RL.
title Upside-Down Reinforcement Learning for More Interpretable Optimal Control
topic Machine Learning
url https://arxiv.org/abs/2411.11457