VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Shen, Haozhan, Liu, Peng, Li, Jingcheng, Fang, Chunxin, Ma, Yibo, Liao, Jiajia, Shen, Qiaoli, Zhang, Zilun, Zhao, Kangjia, Zhang, Qianqian, Xu, Ruochen, Zhao, Tiancheng
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913790655725568
author Shen, Haozhan
Liu, Peng
Li, Jingcheng
Fang, Chunxin
Ma, Yibo
Liao, Jiajia
Shen, Qiaoli
Zhang, Zilun
Zhao, Kangjia
Zhang, Qianqian
Xu, Ruochen
Zhao, Tiancheng
author_facet Shen, Haozhan
Liu, Peng
Li, Jingcheng
Fang, Chunxin
Ma, Yibo
Liao, Jiajia
Shen, Qiaoli
Zhang, Zilun
Zhao, Kangjia
Zhang, Qianqian
Xu, Ruochen
Zhao, Tiancheng
contents Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Motivated by this observation, we investigate the extension of R1-style reinforcement learning to Vision-Language Models (VLMs), aiming to enhance their visual reasoning capabilities. To this end, we develop VLM-R1, a dedicated framework designed to harness RL for improving VLMs' performance on general vision-language tasks. Using this framework, we further explore the feasibility of applying RL to visual domain. Experimental results indicate that the RL-based model not only delivers competitive performance on visual understanding tasks but also surpasses Supervised Fine-Tuning (SFT) in generalization ability. Furthermore, we conduct comprehensive ablation studies that uncover a series of noteworthy insights, including the presence of reward hacking in object detection, the emergence of the "OD aha moment", the impact of training data quality, and the scaling behavior of RL across different model sizes. Through these analyses, we aim to deepen the understanding of how reinforcement learning enhances the capabilities of vision-language models, and we hope our findings and open-source contributions will support continued progress in the vision-language RL community. Our code and model are available at https://github.com/om-ai-lab/VLM-R1
format Preprint
id arxiv_https___arxiv_org_abs_2504_07615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Shen, Haozhan
Liu, Peng
Li, Jingcheng
Fang, Chunxin
Ma, Yibo
Liao, Jiajia
Shen, Qiaoli
Zhang, Zilun
Zhao, Kangjia
Zhang, Qianqian
Xu, Ruochen
Zhao, Tiancheng
Computer Vision and Pattern Recognition
Computation and Language
Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-based reward formulation, which leverages tasks with deterministic ground-truth answers to enable precise and stable reward computation. In the visual domain, we similarly observe that a wide range of visual understanding tasks are inherently equipped with well-defined ground-truth annotations. This property makes them naturally compatible with rule-based reward mechanisms. Motivated by this observation, we investigate the extension of R1-style reinforcement learning to Vision-Language Models (VLMs), aiming to enhance their visual reasoning capabilities. To this end, we develop VLM-R1, a dedicated framework designed to harness RL for improving VLMs' performance on general vision-language tasks. Using this framework, we further explore the feasibility of applying RL to visual domain. Experimental results indicate that the RL-based model not only delivers competitive performance on visual understanding tasks but also surpasses Supervised Fine-Tuning (SFT) in generalization ability. Furthermore, we conduct comprehensive ablation studies that uncover a series of noteworthy insights, including the presence of reward hacking in object detection, the emergence of the "OD aha moment", the impact of training data quality, and the scaling behavior of RL across different model sizes. Through these analyses, we aim to deepen the understanding of how reinforcement learning enhances the capabilities of vision-language models, and we hope our findings and open-source contributions will support continued progress in the vision-language RL community. Our code and model are available at https://github.com/om-ai-lab/VLM-R1
title VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.07615