REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chakraborty, Souradip, Singh, Anukriti, Bhaskar, Amisha, Tokekar, Pratap, Manocha, Dinesh, Bedi, Amrit Singh
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915108410621952
author Chakraborty, Souradip
Singh, Anukriti
Bhaskar, Amisha
Tokekar, Pratap
Manocha, Dinesh
Bedi, Amrit Singh
author_facet Chakraborty, Souradip
Singh, Anukriti
Bhaskar, Amisha
Tokekar, Pratap
Manocha, Dinesh
Bedi, Amrit Singh
contents The effectiveness of reinforcement learning (RL) agents in continuous control robotics tasks is mainly dependent on the design of the underlying reward function, which is highly prone to reward hacking. A misalignment between the reward function and underlying human preferences (values, social norms) can lead to catastrophic outcomes in the real world especially in the context of robotics for critical decision making. Recent methods aim to mitigate misalignment by learning reward functions from human preferences and subsequently performing policy optimization. However, these methods inadvertently introduce a distribution shift during reward learning due to ignoring the dependence of agent-generated trajectories on the reward learning objective, ultimately resulting in sub-optimal alignment. Hence, in this work, we address this challenge by advocating for the adoption of regularized reward functions that more accurately mirror the intended behaviors of the agent. We propose a novel concept of reward regularization within the robotic RLHF (RL from Human Feedback) framework, which we refer to as \emph{agent preferences}. Our approach uniquely incorporates not just human feedback in the form of preferences but also considers the preferences of the RL agent itself during the reward function learning process. This dual consideration significantly mitigates the issue of distribution shift in RLHF with a computationally tractable algorithm. We provide a theoretical justification for the proposed algorithm by formulating the robotic RLHF problem as a bilevel optimization problem and developing a computationally tractable version of the same. We demonstrate the efficiency of our algorithm {\ours} in several continuous control benchmarks in DeepMind Control Suite \cite{tassa2018deepmind}.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14436
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
Chakraborty, Souradip
Singh, Anukriti
Bhaskar, Amisha
Tokekar, Pratap
Manocha, Dinesh
Bedi, Amrit Singh
Robotics
Machine Learning
The effectiveness of reinforcement learning (RL) agents in continuous control robotics tasks is mainly dependent on the design of the underlying reward function, which is highly prone to reward hacking. A misalignment between the reward function and underlying human preferences (values, social norms) can lead to catastrophic outcomes in the real world especially in the context of robotics for critical decision making. Recent methods aim to mitigate misalignment by learning reward functions from human preferences and subsequently performing policy optimization. However, these methods inadvertently introduce a distribution shift during reward learning due to ignoring the dependence of agent-generated trajectories on the reward learning objective, ultimately resulting in sub-optimal alignment. Hence, in this work, we address this challenge by advocating for the adoption of regularized reward functions that more accurately mirror the intended behaviors of the agent. We propose a novel concept of reward regularization within the robotic RLHF (RL from Human Feedback) framework, which we refer to as \emph{agent preferences}. Our approach uniquely incorporates not just human feedback in the form of preferences but also considers the preferences of the RL agent itself during the reward function learning process. This dual consideration significantly mitigates the issue of distribution shift in RLHF with a computationally tractable algorithm. We provide a theoretical justification for the proposed algorithm by formulating the robotic RLHF problem as a bilevel optimization problem and developing a computationally tractable version of the same. We demonstrate the efficiency of our algorithm {\ours} in several continuous control benchmarks in DeepMind Control Suite \cite{tassa2018deepmind}.
title REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
topic Robotics
Machine Learning
url https://arxiv.org/abs/2312.14436