Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Wenxuan, Jia, Bohan, Zhai, Zijie, Cao, Shaosheng, Ye, Zheyu, Zhao, Fei, Xu, Zhe, Tang, Xu, Hu, Yao, Lin, Shaohui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917300878180352
author Huang, Wenxuan
Jia, Bohan
Zhai, Zijie
Cao, Shaosheng
Ye, Zheyu
Zhao, Fei
Xu, Zhe
Tang, Xu
Hu, Yao
Lin, Shaohui
author_facet Huang, Wenxuan
Jia, Bohan
Zhai, Zijie
Cao, Shaosheng
Ye, Zheyu
Zhao, Fei
Xu, Zhe
Tang, Xu
Hu, Yao
Lin, Shaohui
contents DeepSeek-R1-Zero has successfully demonstrated the emergence of reasoning capabilities in LLMs purely through Reinforcement Learning (RL). Inspired by this breakthrough, we explore how RL can be utilized to enhance the reasoning capability of MLLMs. However, direct training with RL struggles to activate complex reasoning capabilities such as questioning and reflection in MLLMs, due to the absence of substantial high-quality multimodal reasoning data. To address this issue, we propose the reasoning MLLM, Vision-R1, to improve multimodal reasoning capability. Specifically, we first construct a high-quality multimodal CoT dataset without human annotations by leveraging an existing MLLM and DeepSeek-R1 through modality bridging and data filtering to obtain a 200K multimodal CoT dataset, Vision-R1-cold dataset. It serves as cold-start initialization data for Vision-R1. To mitigate the optimization challenges caused by overthinking after cold start, we propose Progressive Thinking Suppression Training (PTST) strategy and employ Group Relative Policy Optimization (GRPO) with the hard formatting result reward function to gradually refine the model's ability to learn correct and complex reasoning processes on a 10K multimodal math dataset. Comprehensive experiments show our model achieves an average improvement of $\sim$6% across various multimodal math reasoning benchmarks. Vision-R1-7B achieves a 73.5% accuracy on the widely used MathVista benchmark, which is only 0.4% lower than the leading reasoning model, OpenAI O1. Scaling up the amount of multimodal math data in the RL training, Vision-R1-32B and Vison-R1-72B achieves 76.4% and 78.2% MathVista benchmark scores, respectively. The datasets and code will be released in: https://github.com/Osilly/Vision-R1 .
format Preprint
id arxiv_https___arxiv_org_abs_2503_06749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Huang, Wenxuan
Jia, Bohan
Zhai, Zijie
Cao, Shaosheng
Ye, Zheyu
Zhao, Fei
Xu, Zhe
Tang, Xu
Hu, Yao
Lin, Shaohui
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
DeepSeek-R1-Zero has successfully demonstrated the emergence of reasoning capabilities in LLMs purely through Reinforcement Learning (RL). Inspired by this breakthrough, we explore how RL can be utilized to enhance the reasoning capability of MLLMs. However, direct training with RL struggles to activate complex reasoning capabilities such as questioning and reflection in MLLMs, due to the absence of substantial high-quality multimodal reasoning data. To address this issue, we propose the reasoning MLLM, Vision-R1, to improve multimodal reasoning capability. Specifically, we first construct a high-quality multimodal CoT dataset without human annotations by leveraging an existing MLLM and DeepSeek-R1 through modality bridging and data filtering to obtain a 200K multimodal CoT dataset, Vision-R1-cold dataset. It serves as cold-start initialization data for Vision-R1. To mitigate the optimization challenges caused by overthinking after cold start, we propose Progressive Thinking Suppression Training (PTST) strategy and employ Group Relative Policy Optimization (GRPO) with the hard formatting result reward function to gradually refine the model's ability to learn correct and complex reasoning processes on a 10K multimodal math dataset. Comprehensive experiments show our model achieves an average improvement of $\sim$6% across various multimodal math reasoning benchmarks. Vision-R1-7B achieves a 73.5% accuracy on the widely used MathVista benchmark, which is only 0.4% lower than the leading reasoning model, OpenAI O1. Scaling up the amount of multimodal math data in the RL training, Vision-R1-32B and Vison-R1-72B achieves 76.4% and 78.2% MathVista benchmark scores, respectively. The datasets and code will be released in: https://github.com/Osilly/Vision-R1 .
title Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2503.06749