Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Venkataraman, Sreyas, Wang, Yufei, Wang, Ziyu, Ravie, Navin Sriram, Erickson, Zackory, Held, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909724292677632
author Venkataraman, Sreyas
Wang, Yufei
Wang, Ziyu
Ravie, Navin Sriram
Erickson, Zackory
Held, David
author_facet Venkataraman, Sreyas
Wang, Yufei
Wang, Ziyu
Ravie, Navin Sriram
Erickson, Zackory
Held, David
contents Offline reinforcement learning can enable policy learning from pre-collected, sub-optimal datasets without online interactions. This makes it ideal for real-world robots and safety-critical scenarios, where collecting online data or expert demonstrations is slow, costly, and risky. However, most existing offline RL works assume the dataset is already labeled with the task rewards, a process that often requires significant human effort, especially when ground-truth states are hard to ascertain (e.g., in the real-world). In this paper, we build on prior work, specifically RL-VLM-F, and propose a novel system that automatically generates reward labels for offline datasets using preference feedback from a vision-language model and a text description of the task. Our method then learns a policy using offline RL with the reward-labeled dataset. We demonstrate the system's applicability to a complex real-world robot-assisted dressing task, where we first learn a reward function using a vision-language model on a sub-optimal offline dataset, and then we use the learned reward to employ Implicit Q learning to develop an effective dressing policy. Our method also performs well in simulation tasks involving the manipulation of rigid and deformable objects, and significantly outperform baselines such as behavior cloning and inverse RL. In summary, we propose a new system that enables automatic reward labeling and policy learning from unlabeled, sub-optimal offline datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05273
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Real-World Offline Reinforcement Learning from Vision Language Model Feedback
Venkataraman, Sreyas
Wang, Yufei
Wang, Ziyu
Ravie, Navin Sriram
Erickson, Zackory
Held, David
Robotics
Artificial Intelligence
Machine Learning
Offline reinforcement learning can enable policy learning from pre-collected, sub-optimal datasets without online interactions. This makes it ideal for real-world robots and safety-critical scenarios, where collecting online data or expert demonstrations is slow, costly, and risky. However, most existing offline RL works assume the dataset is already labeled with the task rewards, a process that often requires significant human effort, especially when ground-truth states are hard to ascertain (e.g., in the real-world). In this paper, we build on prior work, specifically RL-VLM-F, and propose a novel system that automatically generates reward labels for offline datasets using preference feedback from a vision-language model and a text description of the task. Our method then learns a policy using offline RL with the reward-labeled dataset. We demonstrate the system's applicability to a complex real-world robot-assisted dressing task, where we first learn a reward function using a vision-language model on a sub-optimal offline dataset, and then we use the learned reward to employ Implicit Q learning to develop an effective dressing policy. Our method also performs well in simulation tasks involving the manipulation of rigid and deformable objects, and significantly outperform baselines such as behavior cloning and inverse RL. In summary, we propose a new system that enables automatic reward labeling and policy learning from unlabeled, sub-optimal offline datasets.
title Real-World Offline Reinforcement Learning from Vision Language Model Feedback
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2411.05273