Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pala, Tej Deep, Sharma, Panshul, Zadeh, Amir, Li, Chuan, Poria, Soujanya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918034782814208
author Pala, Tej Deep
Sharma, Panshul
Zadeh, Amir
Li, Chuan
Poria, Soujanya
author_facet Pala, Tej Deep
Sharma, Panshul
Zadeh, Amir
Li, Chuan
Poria, Soujanya
contents Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Reward Models (PRMs) score each intermediate step to steer generation toward coherent solutions. We introduce PathFinder-PRM, a novel hierarchical, error-aware discriminative PRM that first classifies math and consistency errors at each step, then combines these fine-grained signals to estimate step correctness. To train PathFinder-PRM, we construct a 400K-sample dataset by enriching the human-annotated PRM800K corpus and RLHFlow Mistral traces with three-dimensional step-level labels. On PRMBench, PathFinder-PRM achieves a new state-of-the-art PRMScore of 67.7, outperforming the prior best (65.5) while using 3 times less data. When applied to reward guided greedy search, our model yields prm@8 48.3, a +1.5 point gain over the strongest baseline. These results demonstrate that decoupled error detection and reward estimation not only boost fine-grained error detection but also substantially improve end-to-end, reward-guided mathematical reasoning with greater data efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
Pala, Tej Deep
Sharma, Panshul
Zadeh, Amir
Li, Chuan
Poria, Soujanya
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving. While Outcome Reward Models verify only final answers, Process Reward Models (PRMs) score each intermediate step to steer generation toward coherent solutions. We introduce PathFinder-PRM, a novel hierarchical, error-aware discriminative PRM that first classifies math and consistency errors at each step, then combines these fine-grained signals to estimate step correctness. To train PathFinder-PRM, we construct a 400K-sample dataset by enriching the human-annotated PRM800K corpus and RLHFlow Mistral traces with three-dimensional step-level labels. On PRMBench, PathFinder-PRM achieves a new state-of-the-art PRMScore of 67.7, outperforming the prior best (65.5) while using 3 times less data. When applied to reward guided greedy search, our model yields prm@8 48.3, a +1.5 point gain over the strongest baseline. These results demonstrate that decoupled error detection and reward estimation not only boost fine-grained error detection but also substantially improve end-to-end, reward-guided mathematical reasoning with greater data efficiency.
title Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.19706