Multi Task Inverse Reinforcement Learning for Common Sense Reward

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Glazer, Neta, Navon, Aviv, Shamsian, Aviv, Fetaya, Ethan
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915570219220992
author Glazer, Neta
Navon, Aviv
Shamsian, Aviv
Fetaya, Ethan
author_facet Glazer, Neta
Navon, Aviv
Shamsian, Aviv
Fetaya, Ethan
contents One of the challenges in applying reinforcement learning in a complex real-world environment lies in providing the agent with a sufficiently detailed reward function. Any misalignment between the reward and the desired behavior can result in unwanted outcomes. This may lead to issues like "reward hacking" where the agent maximizes rewards by unintended behavior. In this work, we propose to disentangle the reward into two distinct parts. A simple task-specific reward, outlining the particulars of the task at hand, and an unknown common-sense reward, indicating the expected behavior of the agent within the environment. We then explore how this common-sense reward can be learned from expert demonstrations. We first show that inverse reinforcement learning, even when it succeeds in training an agent, does not learn a useful reward function. That is, training a new agent with the learned reward does not impair the desired behaviors. We then demonstrate that this problem can be solved by training simultaneously on multiple tasks. That is, multi-task inverse reinforcement learning can be applied to learn a useful reward function.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11367
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi Task Inverse Reinforcement Learning for Common Sense Reward
Glazer, Neta
Navon, Aviv
Shamsian, Aviv
Fetaya, Ethan
Machine Learning
One of the challenges in applying reinforcement learning in a complex real-world environment lies in providing the agent with a sufficiently detailed reward function. Any misalignment between the reward and the desired behavior can result in unwanted outcomes. This may lead to issues like "reward hacking" where the agent maximizes rewards by unintended behavior. In this work, we propose to disentangle the reward into two distinct parts. A simple task-specific reward, outlining the particulars of the task at hand, and an unknown common-sense reward, indicating the expected behavior of the agent within the environment. We then explore how this common-sense reward can be learned from expert demonstrations. We first show that inverse reinforcement learning, even when it succeeds in training an agent, does not learn a useful reward function. That is, training a new agent with the learned reward does not impair the desired behaviors. We then demonstrate that this problem can be solved by training simultaneously on multiple tasks. That is, multi-task inverse reinforcement learning can be applied to learn a useful reward function.
title Multi Task Inverse Reinforcement Learning for Common Sense Reward
topic Machine Learning
url https://arxiv.org/abs/2402.11367