A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Congmin, Zhu, Jiachen, Ou, Zhuoying, Chen, Yuxiang, Zhang, Kangning, Shan, Rong, Zheng, Zeyu, Yang, Mengyue, Lin, Jianghao, Yu, Yong, Zhang, Weinan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917444372660224
author Zheng, Congmin
Zhu, Jiachen
Ou, Zhuoying
Chen, Yuxiang
Zhang, Kangning
Shan, Rong
Zheng, Zeyu
Yang, Mengyue
Lin, Jianghao
Yu, Yong
Zhang, Weinan
author_facet Zheng, Congmin
Zhu, Jiachen
Ou, Zhuoying
Chen, Yuxiang
Zhang, Kangning
Shan, Rong
Zheng, Zeyu
Yang, Mengyue
Lin, Jianghao
Yu, Yong
Zhang, Weinan
contents Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) address this gap by evaluating and guiding reasoning at the step or trajectory level. This survey provides a systematic overview of PRMs through the full loop: how to generate process data, build PRMs, and use PRMs for test-time scaling and reinforcement learning. We summarize applications across math, code, text, multimodal reasoning, robotics, and agents, and review emerging benchmarks. Our goal is to clarify design spaces, reveal open challenges, and guide future research toward fine-grained, robust reasoning alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08049
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Zheng, Congmin
Zhu, Jiachen
Ou, Zhuoying
Chen, Yuxiang
Zhang, Kangning
Shan, Rong
Zheng, Zeyu
Yang, Mengyue
Lin, Jianghao
Yu, Yong
Zhang, Weinan
Computation and Language
Artificial Intelligence
Although Large Language Models (LLMs) exhibit advanced reasoning ability, conventional alignment remains largely dominated by outcome reward models (ORMs) that judge only final answers. Process Reward Models(PRMs) address this gap by evaluating and guiding reasoning at the step or trajectory level. This survey provides a systematic overview of PRMs through the full loop: how to generate process data, build PRMs, and use PRMs for test-time scaling and reinforcement learning. We summarize applications across math, code, text, multimodal reasoning, robotics, and agents, and review emerging benchmarks. Our goal is to clarify design spaces, reveal open challenges, and guide future research toward fine-grained, robust reasoning alignment.
title A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.08049