All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Xinyu, Zou, Shu, Yang, Zhaoyuan, He, Mengqi, Tu, Peter, Zhang, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908930214461440
author Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Tu, Peter
Zhang, Jing
author_facet Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Tu, Peter
Zhang, Jing
contents Recent studies have demonstrated that Reinforcement Learning (RL), notably Group Relative Policy Optimization (GRPO), can intrinsically elicit and enhance the reasoning capabilities of Vision-Language Models (VLMs). However, despite the promise, the underlying mechanisms that drive the effectiveness of RL models as well as their limitations remain underexplored. In this paper, we highlight a fundamental behavioral distinction between RL and base models, where the former engages in deeper yet narrow reasoning, while base models, despite less refined along individual path, exhibit broader and more diverse thinking patterns. Through further analysis of training dynamics, we show that GRPO is prone to diversity collapse, causing models to prematurely converge to a limited subset of reasoning strategies while discarding the majority of potential alternatives, leading to local optima and poor scalability. To address this, we propose Multi-Group Policy Optimization (MUPO), a simple yet effective approach designed to incentivize divergent thinking across multiple solutions, and demonstrate its effectiveness on established benchmarks. Project page: https://xytian1008.github.io/MUPO/
format Preprint
id arxiv_https___arxiv_org_abs_2604_00479
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
Tian, Xinyu
Zou, Shu
Yang, Zhaoyuan
He, Mengqi
Tu, Peter
Zhang, Jing
Computer Vision and Pattern Recognition
Recent studies have demonstrated that Reinforcement Learning (RL), notably Group Relative Policy Optimization (GRPO), can intrinsically elicit and enhance the reasoning capabilities of Vision-Language Models (VLMs). However, despite the promise, the underlying mechanisms that drive the effectiveness of RL models as well as their limitations remain underexplored. In this paper, we highlight a fundamental behavioral distinction between RL and base models, where the former engages in deeper yet narrow reasoning, while base models, despite less refined along individual path, exhibit broader and more diverse thinking patterns. Through further analysis of training dynamics, we show that GRPO is prone to diversity collapse, causing models to prematurely converge to a limited subset of reasoning strategies while discarding the majority of potential alternatives, leading to local optima and poor scalability. To address this, we propose Multi-Group Policy Optimization (MUPO), a simple yet effective approach designed to incentivize divergent thinking across multiple solutions, and demonstrate its effectiveness on established benchmarks. Project page: https://xytian1008.github.io/MUPO/
title All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.00479