Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jiachen, Zheng, Congmin, Lin, Jianghao, Du, Kounianhua, Wen, Ying, Yu, Yong, Wang, Jun, Zhang, Weinan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917930749394944
author Zhu, Jiachen
Zheng, Congmin
Lin, Jianghao
Du, Kounianhua
Wen, Ying
Yu, Yong
Wang, Jun
Zhang, Weinan
author_facet Zhu, Jiachen
Zheng, Congmin
Lin, Jianghao
Du, Kounianhua
Wen, Ying
Yu, Yong
Wang, Jun
Zhang, Weinan
contents While large language models (LLMs) have significantly advanced mathematical reasoning, Process Reward Models (PRMs) have been developed to evaluate the logical validity of reasoning steps. However, PRMs still struggle with out-of-distribution (OOD) challenges. This paper identifies key OOD issues, including step OOD, caused by differences in reasoning patterns across model types and sizes, and question OOD, which arises from dataset shifts between training data and real-world problems. To address these issues, we introduce Retrieval-Augmented Process Reward Model (RetrievalPRM), a novel framework designed to tackle these OOD issues. By utilizing a two-stage retrieval-enhanced mechanism, RetrievalPRM retrieves semantically similar questions and steps as a warmup, enhancing PRM's ability to evaluate target steps and improving generalization and reasoning consistency across different models and problem types. Our extensive experiments demonstrate that RetrievalPRM outperforms existing baselines across multiple real-world datasets. Our open-source contributions include a retrieval-enhanced dataset, a tuning framework for PRM training, and the RetrievalPRM model, establishing a new standard for PRM performance.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
Zhu, Jiachen
Zheng, Congmin
Lin, Jianghao
Du, Kounianhua
Wen, Ying
Yu, Yong
Wang, Jun
Zhang, Weinan
Artificial Intelligence
Information Retrieval
While large language models (LLMs) have significantly advanced mathematical reasoning, Process Reward Models (PRMs) have been developed to evaluate the logical validity of reasoning steps. However, PRMs still struggle with out-of-distribution (OOD) challenges. This paper identifies key OOD issues, including step OOD, caused by differences in reasoning patterns across model types and sizes, and question OOD, which arises from dataset shifts between training data and real-world problems. To address these issues, we introduce Retrieval-Augmented Process Reward Model (RetrievalPRM), a novel framework designed to tackle these OOD issues. By utilizing a two-stage retrieval-enhanced mechanism, RetrievalPRM retrieves semantically similar questions and steps as a warmup, enhancing PRM's ability to evaluate target steps and improving generalization and reasoning consistency across different models and problem types. Our extensive experiments demonstrate that RetrievalPRM outperforms existing baselines across multiple real-world datasets. Our open-source contributions include a retrieval-enhanced dataset, a tuning framework for PRM training, and the RetrievalPRM model, establishing a new standard for PRM performance.
title Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning
topic Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2502.14361