Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choe, Jean Seong Bjorn, Choi, Bumkyu, Kim, Jong-kook
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914948779606016
author Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
author_facet Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
contents This report presents a solution for the swing-up and stabilisation tasks of the acrobot and the pendubot, developed for the AI Olympics competition at IROS 2024. Our approach employs the Average-Reward Entropy Advantage Policy Optimization (AR-EAPO), a model-free reinforcement learning (RL) algorithm that combines average-reward RL and maximum entropy RL. Results demonstrate that our controller achieves improved performance and robustness scores compared to established baseline methods in both the acrobot and pendubot scenarios, without the need for a heavily engineered reward function or system model. The current results are applicable exclusively to the simulation stage setup.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08938
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
Choe, Jean Seong Bjorn
Choi, Bumkyu
Kim, Jong-kook
Robotics
Machine Learning
This report presents a solution for the swing-up and stabilisation tasks of the acrobot and the pendubot, developed for the AI Olympics competition at IROS 2024. Our approach employs the Average-Reward Entropy Advantage Policy Optimization (AR-EAPO), a model-free reinforcement learning (RL) algorithm that combines average-reward RL and maximum entropy RL. Results demonstrate that our controller achieves improved performance and robustness scores compared to established baseline methods in both the acrobot and pendubot scenarios, without the need for a heavily engineered reward function or system model. The current results are applicable exclusively to the simulation stage setup.
title Average-Reward Maximum Entropy Reinforcement Learning for Underactuated Double Pendulum Tasks
topic Robotics
Machine Learning
url https://arxiv.org/abs/2409.08938