Better Process Supervision with Bi-directional Rewarding Signals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Wenxiang, He, Wei, Xi, Zhiheng, Guo, Honglin, Hong, Boyang, Zhang, Jiazheng, Zheng, Rui, Li, Nijun, Gui, Tao, Li, Yun, Zhang, Qi, Huang, Xuanjing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917947182678016
author Chen, Wenxiang
He, Wei
Xi, Zhiheng
Guo, Honglin
Hong, Boyang
Zhang, Jiazheng
Zheng, Rui
Li, Nijun
Gui, Tao
Li, Yun
Zhang, Qi
Huang, Xuanjing
author_facet Chen, Wenxiang
He, Wei
Xi, Zhiheng
Guo, Honglin
Hong, Boyang
Zhang, Jiazheng
Zheng, Rui
Li, Nijun
Gui, Tao
Li, Yun
Zhang, Qi
Huang, Xuanjing
contents Process supervision, i.e., evaluating each step, is critical for complex large language model (LLM) reasoning and test-time searching with increased inference compute. Existing approaches, represented by process reward models (PRMs), primarily focus on rewarding signals up to the current step, exhibiting a one-directional nature and lacking a mechanism to model the distance to the final target. To address this problem, we draw inspiration from the A* algorithm, which states that an effective supervisory signal should simultaneously consider the incurred cost and the estimated cost for reaching the target. Building on this key insight, we introduce BiRM, a novel process supervision model that not only evaluates the correctness of previous steps but also models the probability of future success. We conduct extensive experiments on mathematical reasoning tasks and demonstrate that BiRM provides more precise evaluations of LLM reasoning steps, achieving an improvement of 3.1% on Gaokao2023 over PRM under the Best-of-N sampling method. Besides, in search-based strategies, BiRM provides more comprehensive guidance and outperforms ORM by 5.0% and PRM by 3.8% respectively on MATH-500.
format Preprint
id arxiv_https___arxiv_org_abs_2503_04618
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Better Process Supervision with Bi-directional Rewarding Signals
Chen, Wenxiang
He, Wei
Xi, Zhiheng
Guo, Honglin
Hong, Boyang
Zhang, Jiazheng
Zheng, Rui
Li, Nijun
Gui, Tao
Li, Yun
Zhang, Qi
Huang, Xuanjing
Computation and Language
Process supervision, i.e., evaluating each step, is critical for complex large language model (LLM) reasoning and test-time searching with increased inference compute. Existing approaches, represented by process reward models (PRMs), primarily focus on rewarding signals up to the current step, exhibiting a one-directional nature and lacking a mechanism to model the distance to the final target. To address this problem, we draw inspiration from the A* algorithm, which states that an effective supervisory signal should simultaneously consider the incurred cost and the estimated cost for reaching the target. Building on this key insight, we introduce BiRM, a novel process supervision model that not only evaluates the correctness of previous steps but also models the probability of future success. We conduct extensive experiments on mathematical reasoning tasks and demonstrate that BiRM provides more precise evaluations of LLM reasoning steps, achieving an improvement of 3.1% on Gaokao2023 over PRM under the Best-of-N sampling method. Besides, in search-based strategies, BiRM provides more comprehensive guidance and outperforms ORM by 5.0% and PRM by 3.8% respectively on MATH-500.
title Better Process Supervision with Bi-directional Rewarding Signals
topic Computation and Language
url https://arxiv.org/abs/2503.04618