Training Data Efficiency in Multimodal Process Reward Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Jinyuan, Huang, Chengsong, Huang, Langlin, Xu, Shaoyang, Liu, Haolin, Zhang, Wenxuan, Huang, Jiaxin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917248595132416
author Li, Jinyuan
Huang, Chengsong
Huang, Langlin
Xu, Shaoyang
Liu, Haolin
Zhang, Wenxuan
Huang, Jiaxin
author_facet Li, Jinyuan
Huang, Chengsong
Huang, Langlin
Xu, Shaoyang
Liu, Haolin
Zhang, Wenxuan
Huang, Jiaxin
contents Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)-annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training. Our preliminary experiments reveal that MPRM training quickly saturates under random subsampling of the training data, indicating substantial redundancy within existing MC-annotated corpora. To explain this, we formalize a theoretical framework and reveal that informative gradient updates depend on two factors: label mixtures of positive/negative steps and label reliability (average MC scores of positive steps). Guided by these insights, we propose the Balanced-Information Score (BIS), which prioritizes both mixture and reliability based on existing MC signals at the rollout level, without incurring any additional cost. Across two backbones (InternVL2.5-8B and Qwen2.5-VL-7B) on VisualProcessBench, BIS-selected subsets consistently match and even surpass the full-data performance at small fractions. Notably, the BIS subset reaches full-data performance using only 10% of the training data, improving over random subsampling by a relative 4.1%.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04145
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training Data Efficiency in Multimodal Process Reward Models
Li, Jinyuan
Huang, Chengsong
Huang, Langlin
Xu, Shaoyang
Liu, Haolin
Zhang, Wenxuan
Huang, Jiaxin
Machine Learning
Computation and Language
Multimedia
Multimodal Process Reward Models (MPRMs) are central to step-level supervision for visual reasoning in MLLMs. Training MPRMs typically requires large-scale Monte Carlo (MC)-annotated corpora, incurring substantial training cost. This paper studies the data efficiency for MPRM training. Our preliminary experiments reveal that MPRM training quickly saturates under random subsampling of the training data, indicating substantial redundancy within existing MC-annotated corpora. To explain this, we formalize a theoretical framework and reveal that informative gradient updates depend on two factors: label mixtures of positive/negative steps and label reliability (average MC scores of positive steps). Guided by these insights, we propose the Balanced-Information Score (BIS), which prioritizes both mixture and reliability based on existing MC signals at the rollout level, without incurring any additional cost. Across two backbones (InternVL2.5-8B and Qwen2.5-VL-7B) on VisualProcessBench, BIS-selected subsets consistently match and even surpass the full-data performance at small fractions. Notably, the BIS subset reaches full-data performance using only 10% of the training data, improving over random subsampling by a relative 4.1%.
title Training Data Efficiency in Multimodal Process Reward Models
topic Machine Learning
Computation and Language
Multimedia
url https://arxiv.org/abs/2602.04145