R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Huanjin, Yin, Qixiang, Zhang, Jingyi, Yang, Min, Wang, Yibo, Wu, Wenhao, Su, Fei, Shen, Li, Qiu, Minghui, Tao, Dacheng, Huang, Jiaxing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910962672467968
author Yao, Huanjin
Yin, Qixiang
Zhang, Jingyi
Yang, Min
Wang, Yibo
Wu, Wenhao
Su, Fei
Shen, Li
Qiu, Minghui
Tao, Dacheng
Huang, Jiaxing
author_facet Yao, Huanjin
Yin, Qixiang
Zhang, Jingyi
Yang, Min
Wang, Yibo
Wu, Wenhao
Su, Fei
Shen, Li
Qiu, Minghui
Tao, Dacheng
Huang, Jiaxing
contents In this work, we aim to incentivize the reasoning ability of Multimodal Large Language Models (MLLMs) via reinforcement learning (RL) and develop an effective approach that mitigates the sparse reward and advantage vanishing issues during RL. To this end, we propose Share-GRPO, a novel RL approach that tackle these issues by exploring and sharing diverse reasoning trajectories over expanded question space. Specifically, Share-GRPO first expands the question space for a given question via data transformation techniques, and then encourages MLLM to effectively explore diverse reasoning trajectories over the expanded question space and shares the discovered reasoning trajectories across the expanded questions during RL. In addition, Share-GRPO also shares reward information during advantage computation, which estimates solution advantages hierarchically across and within question variants, allowing more accurate estimation of relative advantages and improving the stability of policy training. Extensive evaluations over six widely-used reasoning benchmarks showcase the superior performance of our method. Code will be available at https://github.com/HJYao00/R1-ShareVL.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16673
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
Yao, Huanjin
Yin, Qixiang
Zhang, Jingyi
Yang, Min
Wang, Yibo
Wu, Wenhao
Su, Fei
Shen, Li
Qiu, Minghui
Tao, Dacheng
Huang, Jiaxing
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
In this work, we aim to incentivize the reasoning ability of Multimodal Large Language Models (MLLMs) via reinforcement learning (RL) and develop an effective approach that mitigates the sparse reward and advantage vanishing issues during RL. To this end, we propose Share-GRPO, a novel RL approach that tackle these issues by exploring and sharing diverse reasoning trajectories over expanded question space. Specifically, Share-GRPO first expands the question space for a given question via data transformation techniques, and then encourages MLLM to effectively explore diverse reasoning trajectories over the expanded question space and shares the discovered reasoning trajectories across the expanded questions during RL. In addition, Share-GRPO also shares reward information during advantage computation, which estimates solution advantages hierarchically across and within question variants, allowing more accurate estimation of relative advantages and improving the stability of policy training. Extensive evaluations over six widely-used reasoning benchmarks showcase the superior performance of our method. Code will be available at https://github.com/HJYao00/R1-ShareVL.
title R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.16673