Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Yuexiang, Bai, Hao, Lin, Zipeng, Pan, Jiayi, Tong, Shengbang, Zhou, Yifei, Suhr, Alane, Xie, Saining, LeCun, Yann, Ma, Yi, Levine, Sergey
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917796439392256
author Zhai, Yuexiang
Bai, Hao
Lin, Zipeng
Pan, Jiayi
Tong, Shengbang
Zhou, Yifei
Suhr, Alane
Xie, Saining
LeCun, Yann
Ma, Yi
Levine, Sergey
author_facet Zhai, Yuexiang
Bai, Hao
Lin, Zipeng
Pan, Jiayi
Tong, Shengbang
Zhou, Yifei
Suhr, Alane
Xie, Saining
LeCun, Yann
Ma, Yi
Levine, Sergey
contents Large vision-language models (VLMs) fine-tuned on specialized visual instruction-following data have exhibited impressive language reasoning capabilities across various scenarios. However, this fine-tuning paradigm may not be able to efficiently learn optimal decision-making agents in multi-step goal-directed tasks from interactive environments. To address this challenge, we propose an algorithmic framework that fine-tunes VLMs with reinforcement learning (RL). Specifically, our framework provides a task description and then prompts the VLM to generate chain-of-thought (CoT) reasoning, enabling the VLM to efficiently explore intermediate reasoning steps that lead to the final text-based action. Next, the open-ended text output is parsed into an executable action to interact with the environment to obtain goal-directed task rewards. Finally, our framework uses these task rewards to fine-tune the entire VLM with RL. Empirically, we demonstrate that our proposed framework enhances the decision-making capabilities of VLM agents across various tasks, enabling 7b models to outperform commercial models such as GPT4-V or Gemini. Furthermore, we find that CoT reasoning is a crucial component for performance improvement, as removing the CoT reasoning results in a significant decrease in the overall performance of our method.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10292
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
Zhai, Yuexiang
Bai, Hao
Lin, Zipeng
Pan, Jiayi
Tong, Shengbang
Zhou, Yifei
Suhr, Alane
Xie, Saining
LeCun, Yann
Ma, Yi
Levine, Sergey
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
Large vision-language models (VLMs) fine-tuned on specialized visual instruction-following data have exhibited impressive language reasoning capabilities across various scenarios. However, this fine-tuning paradigm may not be able to efficiently learn optimal decision-making agents in multi-step goal-directed tasks from interactive environments. To address this challenge, we propose an algorithmic framework that fine-tunes VLMs with reinforcement learning (RL). Specifically, our framework provides a task description and then prompts the VLM to generate chain-of-thought (CoT) reasoning, enabling the VLM to efficiently explore intermediate reasoning steps that lead to the final text-based action. Next, the open-ended text output is parsed into an executable action to interact with the environment to obtain goal-directed task rewards. Finally, our framework uses these task rewards to fine-tune the entire VLM with RL. Empirically, we demonstrate that our proposed framework enhances the decision-making capabilities of VLM agents across various tasks, enabling 7b models to outperform commercial models such as GPT4-V or Gemini. Furthermore, we find that CoT reasoning is a crucial component for performance improvement, as removing the CoT reasoning results in a significant decrease in the overall performance of our method.
title Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement Learning
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.10292