Adaptive Reward Design for Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Minjae, ElSayed-Aly, Ingy, Feng, Lu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908368066576384
author Kwon, Minjae
ElSayed-Aly, Ingy
Feng, Lu
author_facet Kwon, Minjae
ElSayed-Aly, Ingy
Feng, Lu
contents There is a surge of interest in using formal languages such as Linear Temporal Logic (LTL) to precisely and succinctly specify complex tasks and derive reward functions for Reinforcement Learning (RL). However, existing methods often assign sparse rewards (e.g., giving a reward of 1 only if a task is completed and 0 otherwise). By providing feedback solely upon task completion, these methods fail to encourage successful subtask completion. This is particularly problematic in environments with inherent uncertainty, where task completion may be unreliable despite progress on intermediate goals. To address this limitation, we propose a suite of reward functions that incentivize an RL agent to complete a task specified by an LTL formula as much as possible, and develop an adaptive reward shaping approach that dynamically updates reward functions during the learning process. Experimental results on a range of benchmark RL environments demonstrate that the proposed approach generally outperforms baselines, achieving earlier convergence to a better policy with higher expected return and task completion rate.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10917
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Reward Design for Reinforcement Learning
Kwon, Minjae
ElSayed-Aly, Ingy
Feng, Lu
Robotics
Artificial Intelligence
Machine Learning
68T40 (Primary) 93E35, 03B44 (Secondary)
There is a surge of interest in using formal languages such as Linear Temporal Logic (LTL) to precisely and succinctly specify complex tasks and derive reward functions for Reinforcement Learning (RL). However, existing methods often assign sparse rewards (e.g., giving a reward of 1 only if a task is completed and 0 otherwise). By providing feedback solely upon task completion, these methods fail to encourage successful subtask completion. This is particularly problematic in environments with inherent uncertainty, where task completion may be unreliable despite progress on intermediate goals. To address this limitation, we propose a suite of reward functions that incentivize an RL agent to complete a task specified by an LTL formula as much as possible, and develop an adaptive reward shaping approach that dynamically updates reward functions during the learning process. Experimental results on a range of benchmark RL environments demonstrate that the proposed approach generally outperforms baselines, achieving earlier convergence to a better policy with higher expected return and task completion rate.
title Adaptive Reward Design for Reinforcement Learning
topic Robotics
Artificial Intelligence
Machine Learning
68T40 (Primary) 93E35, 03B44 (Secondary)
url https://arxiv.org/abs/2412.10917