Code as Reward: Empowering Reinforcement Learning with VLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Venuto, David, Islam, Sami Nur, Klissarov, Martin, Precup, Doina, Yang, Sherry, Anand, Ankit
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910321207148544
author Venuto, David
Islam, Sami Nur
Klissarov, Martin
Precup, Doina
Yang, Sherry
Anand, Ankit
author_facet Venuto, David
Islam, Sami Nur
Klissarov, Martin
Precup, Doina
Yang, Sherry
Anand, Ankit
contents Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In this paper, we aim to leverage these capabilities to support the training of reinforcement learning (RL) agents. In principle, VLMs are well suited for this purpose, as they can naturally analyze image-based observations and provide feedback (reward) on learning progress. However, inference in VLMs is computationally expensive, so querying them frequently to compute rewards would significantly slowdown the training of an RL agent. To address this challenge, we propose a framework named Code as Reward (VLM-CaR). VLM-CaR produces dense reward functions from VLMs through code generation, thereby significantly reducing the computational burden of querying the VLM directly. We show that the dense rewards generated through our approach are very accurate across a diverse set of discrete and continuous environments, and can be more effective in training RL policies than the original sparse environment rewards.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04764
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Code as Reward: Empowering Reinforcement Learning with VLMs
Venuto, David
Islam, Sami Nur
Klissarov, Martin
Precup, Doina
Yang, Sherry
Anand, Ankit
Machine Learning
Pre-trained Vision-Language Models (VLMs) are able to understand visual concepts, describe and decompose complex tasks into sub-tasks, and provide feedback on task completion. In this paper, we aim to leverage these capabilities to support the training of reinforcement learning (RL) agents. In principle, VLMs are well suited for this purpose, as they can naturally analyze image-based observations and provide feedback (reward) on learning progress. However, inference in VLMs is computationally expensive, so querying them frequently to compute rewards would significantly slowdown the training of an RL agent. To address this challenge, we propose a framework named Code as Reward (VLM-CaR). VLM-CaR produces dense reward functions from VLMs through code generation, thereby significantly reducing the computational burden of querying the VLM directly. We show that the dense rewards generated through our approach are very accurate across a diverse set of discrete and continuous environments, and can be more effective in training RL policies than the original sparse environment rewards.
title Code as Reward: Empowering Reinforcement Learning with VLMs
topic Machine Learning
url https://arxiv.org/abs/2402.04764