Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ishihara, Yu, Takasugi, Noriaki, Kawakami, Kotaro, Kinoshita, Masaya, Aoyama, Kazumi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912181445984256
author Ishihara, Yu
Takasugi, Noriaki
Kawakami, Kotaro
Kinoshita, Masaya
Aoyama, Kazumi
author_facet Ishihara, Yu
Takasugi, Noriaki
Kawakami, Kotaro
Kinoshita, Masaya
Aoyama, Kazumi
contents Reinforcement learning has become an essential algorithm for generating complex robotic behaviors. However, to learn such behaviors, it is necessary to design a reward function that describes the task, which often consists of multiple objectives that needs to be balanced. This tuning process is known as reward engineering and typically involves extensive trial-and-error. In this paper, to avoid this trial-and-error process, we propose the concept of Constraints as Rewards (CaR). CaR formulates the task objective using multiple constraint functions instead of a reward function and solves a reinforcement learning problem with constraints using the Lagrangian-method. By adopting this approach, different objectives are automatically balanced, because Lagrange multipliers serves as the weights among the objectives. In addition, we will demonstrate that constraints, expressed as inequalities, provide an intuitive interpretation of the optimization target designed for the task. We apply the proposed method to the standing-up motion generation task of a six-wheeled-telescopic-legged robot and demonstrate that the proposed method successfully acquires the target behavior, even though it is challenging to learn with manually designed reward functions.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04228
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
Ishihara, Yu
Takasugi, Noriaki
Kawakami, Kotaro
Kinoshita, Masaya
Aoyama, Kazumi
Robotics
Artificial Intelligence
Machine Learning
Reinforcement learning has become an essential algorithm for generating complex robotic behaviors. However, to learn such behaviors, it is necessary to design a reward function that describes the task, which often consists of multiple objectives that needs to be balanced. This tuning process is known as reward engineering and typically involves extensive trial-and-error. In this paper, to avoid this trial-and-error process, we propose the concept of Constraints as Rewards (CaR). CaR formulates the task objective using multiple constraint functions instead of a reward function and solves a reinforcement learning problem with constraints using the Lagrangian-method. By adopting this approach, different objectives are automatically balanced, because Lagrange multipliers serves as the weights among the objectives. In addition, we will demonstrate that constraints, expressed as inequalities, provide an intuitive interpretation of the optimization target designed for the task. We apply the proposed method to the standing-up motion generation task of a six-wheeled-telescopic-legged robot and demonstrate that the proposed method successfully acquires the target behavior, even though it is challenging to learn with manually designed reward functions.
title Constraints as Rewards: Reinforcement Learning for Robots without Reward Functions
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.04228