Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choe, Jean Seong Bjorn, Choi, Bumkyu, Kim, Jong-kook
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914948779606016
author Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
author_facet Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
contents This report presents a solution for the swing-up and stabilisation tasks of the acrobot and the pendubot, developed for the AI Olympics competition at IROS 2024. Our approach employs the Average-Reward Entropy Advantage Policy Optimization (AR-EAPO), a model-free reinforcement learning (RL) algorithm that combines average-reward RL and maximum entropy RL. Results demonstrate that our controller achieves improved performance and robustness scores compared to established baseline methods in both the acrobot and pendubot scenarios, without the need for a heavily engineered reward function or system model. The current results are applicable exclusively to the simulation stage setup.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08938
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
Robotics
Machine Learning
This report presents a solution for the swing-up and stabilisation tasks of the acrobot and the pendubot, developed for the AI Olympics competition at IROS 2024. Our approach employs the Average-Reward Entropy Advantage Policy Optimization (AR-EAPO), a model-free reinforcement learning (RL) algorithm that combines average-reward RL and maximum entropy RL. Results demonstrate that our controller achieves improved performance and robustness scores compared to established baseline methods in both the acrobot and pendubot scenarios, without the need for a heavily engineered reward function or system model. The current results are applicable exclusively to the simulation stage setup.
title Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
topic Robotics
Machine Learning
url https://arxiv.org/abs/2409.08938