GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhao, Jian, Liu, Runze, Zhang, Kaiyan, Zhou, Zhimu, Gao, Junqi, Li, Dong, Lyu, Jiafei, Qian, Zhouyi, Qi, Biqing, Li, Xiu, Zhou, Bowen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909566434803712
author Zhao, Jian
Liu, Runze
Zhang, Kaiyan
Zhou, Zhimu
Gao, Junqi
Li, Dong
Lyu, Jiafei
Qian, Zhouyi
Qi, Biqing
Li, Xiu
Zhou, Bowen
author_facet Zhao, Jian
Liu, Runze
Zhang, Kaiyan
Zhou, Zhimu
Gao, Junqi
Li, Dong
Lyu, Jiafei
Qian, Zhouyi
Qi, Biqing
Li, Xiu
Zhou, Bowen
contents Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Our code, model, and data will be available in https://ryanliu112.github.io/GenPRM.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00891
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Zhao, Jian
Liu, Runze
Zhang, Kaiyan
Zhou, Zhimu
Gao, Junqi
Li, Dong
Lyu, Jiafei
Qian, Zhouyi
Qi, Biqing
Li, Xiu
Zhou, Bowen
Computation and Language
Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Our code, model, and data will be available in https://ryanliu112.github.io/GenPRM.
title GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
topic Computation and Language
url https://arxiv.org/abs/2504.00891