VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Weiyun, Gao, Zhangwei, Chen, Lianjie, Chen, Zhe, Zhu, Jinguo, Zhao, Xiangyu, Liu, Yangzhou, Cao, Yue, Ye, Shenglong, Zhu, Xizhou, Lu, Lewei, Duan, Haodong, Qiao, Yu, Dai, Jifeng, Wang, Wenhai
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915196140781568
author Wang, Weiyun
Gao, Zhangwei
Chen, Lianjie
Chen, Zhe
Zhu, Jinguo
Zhao, Xiangyu
Liu, Yangzhou
Cao, Yue
Ye, Shenglong
Zhu, Xizhou
Lu, Lewei
Duan, Haodong
Qiao, Yu
Dai, Jifeng
Wang, Wenhai
author_facet Wang, Weiyun
Gao, Zhangwei
Chen, Lianjie
Chen, Zhe
Zhu, Jinguo
Zhao, Xiangyu
Liu, Yangzhou
Cao, Yue
Ye, Shenglong
Zhu, Xizhou
Lu, Lewei
Duan, Haodong
Qiao, Yu
Dai, Jifeng
Wang, Wenhai
contents We introduce VisualPRM, an advanced multimodal Process Reward Model (PRM) with 8B parameters, which improves the reasoning abilities of existing Multimodal Large Language Models (MLLMs) across different model scales and families with Best-of-N (BoN) evaluation strategies. Specifically, our model improves the reasoning performance of three types of MLLMs and four different model scales. Even when applied to the highly capable InternVL2.5-78B, it achieves a 5.9-point improvement across seven multimodal reasoning benchmarks. Experimental results show that our model exhibits superior performance compared to Outcome Reward Models and Self-Consistency during BoN evaluation. To facilitate the training of multimodal PRMs, we construct a multimodal process supervision dataset VisualPRM400K using an automated data pipeline. For the evaluation of multimodal PRMs, we propose VisualProcessBench, a benchmark with human-annotated step-wise correctness labels, to measure the abilities of PRMs to detect erroneous steps in multimodal reasoning tasks. We hope that our work can inspire more future research and contribute to the development of MLLMs. Our model, data, and benchmark are released in https://internvl.github.io/blog/2025-03-13-VisualPRM/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10291
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Wang, Weiyun
Gao, Zhangwei
Chen, Lianjie
Chen, Zhe
Zhu, Jinguo
Zhao, Xiangyu
Liu, Yangzhou
Cao, Yue
Ye, Shenglong
Zhu, Xizhou
Lu, Lewei
Duan, Haodong
Qiao, Yu
Dai, Jifeng
Wang, Wenhai
Computer Vision and Pattern Recognition
Computation and Language
We introduce VisualPRM, an advanced multimodal Process Reward Model (PRM) with 8B parameters, which improves the reasoning abilities of existing Multimodal Large Language Models (MLLMs) across different model scales and families with Best-of-N (BoN) evaluation strategies. Specifically, our model improves the reasoning performance of three types of MLLMs and four different model scales. Even when applied to the highly capable InternVL2.5-78B, it achieves a 5.9-point improvement across seven multimodal reasoning benchmarks. Experimental results show that our model exhibits superior performance compared to Outcome Reward Models and Self-Consistency during BoN evaluation. To facilitate the training of multimodal PRMs, we construct a multimodal process supervision dataset VisualPRM400K using an automated data pipeline. For the evaluation of multimodal PRMs, we propose VisualProcessBench, a benchmark with human-annotated step-wise correctness labels, to measure the abilities of PRMs to detect erroneous steps in multimodal reasoning tasks. We hope that our work can inspire more future research and contribute to the development of MLLMs. Our model, data, and benchmark are released in https://internvl.github.io/blog/2025-03-13-VisualPRM/.
title VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2503.10291