R-PRM: Reasoning-Driven Process Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: She, Shuaijie, Liu, Junxiao, Liu, Yifeng, Chen, Jiajun, Huang, Xin, Huang, Shujian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916663869308928
author She, Shuaijie
Liu, Junxiao
Liu, Yifeng
Chen, Jiajun
Huang, Xin
Huang, Shujian
author_facet She, Shuaijie
Liu, Junxiao
Liu, Yifeng
Chen, Jiajun
Huang, Xin
Huang, Shujian
contents Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by evaluating each reasoning step. However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy, which is further exacerbated by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM). First, we leverage stronger LLMs to generate seed data from limited annotations, effectively bootstrapping our model's reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we further enhance performance through preference optimization, without requiring additional annotated data. Third, we introduce inference-time scaling to fully harness the model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 11.9 and 8.5 points in F1 scores, respectively. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.5 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and stronger generalization capabilities, thereby highlighting its significant potential.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21295
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle R-PRM: Reasoning-Driven Process Reward Modeling
She, Shuaijie
Liu, Junxiao
Liu, Yifeng
Chen, Jiajun
Huang, Xin
Huang, Shujian
Computation and Language
Large language models (LLMs) inevitably make mistakes when performing step-by-step mathematical reasoning. Process Reward Models (PRMs) have emerged as a promising solution by evaluating each reasoning step. However, existing PRMs typically output evaluation scores directly, limiting both learning efficiency and evaluation accuracy, which is further exacerbated by the scarcity of annotated data. To address these issues, we propose Reasoning-Driven Process Reward Modeling (R-PRM). First, we leverage stronger LLMs to generate seed data from limited annotations, effectively bootstrapping our model's reasoning capabilities and enabling comprehensive step-by-step evaluation. Second, we further enhance performance through preference optimization, without requiring additional annotated data. Third, we introduce inference-time scaling to fully harness the model's reasoning potential. Extensive experiments demonstrate R-PRM's effectiveness: on ProcessBench and PRMBench, it surpasses strong baselines by 11.9 and 8.5 points in F1 scores, respectively. When applied to guide mathematical reasoning, R-PRM achieves consistent accuracy improvements of over 8.5 points across six challenging datasets. Further analysis reveals that R-PRM exhibits more comprehensive evaluation and stronger generalization capabilities, thereby highlighting its significant potential.
title R-PRM: Reasoning-Driven Process Reward Modeling
topic Computation and Language
url https://arxiv.org/abs/2503.21295