From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhengyu, Wang, Yudong, Xiao, Teng, Zhou, Ruochen, Yang, Xuesheng, Wang, Wei, Sui, Zhifang, Wang, Jingang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910977518206976
author Chen, Zhengyu
Wang, Yudong
Xiao, Teng
Zhou, Ruochen
Yang, Xuesheng
Wang, Wei
Sui, Zhifang
Wang, Jingang
author_facet Chen, Zhengyu
Wang, Yudong
Xiao, Teng
Zhou, Ruochen
Yang, Xuesheng
Wang, Wei
Sui, Zhifang
Wang, Jingang
contents Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
Chen, Zhengyu
Wang, Yudong
Xiao, Teng
Zhou, Ruochen
Yang, Xuesheng
Wang, Wei
Sui, Zhifang
Wang, Jingang
Computation and Language
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization.
title From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
topic Computation and Language
url https://arxiv.org/abs/2506.00027