Efficient Process Reward Model Training via Active Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Keyu, Liu, Zichen, Mao, Xin, Pang, Tianyu, Chen, Changyu, Chen, Qiguang, Shieh, Michael Qizhe, Dou, Longxu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913794451570688
author Duan, Keyu
Liu, Zichen
Mao, Xin
Pang, Tianyu
Chen, Changyu
Chen, Qiguang
Shieh, Michael Qizhe
Dou, Longxu
author_facet Duan, Keyu
Liu, Zichen
Mao, Xin
Pang, Tianyu
Chen, Changyu
Chen, Qiguang
Shieh, Michael Qizhe
Dou, Longxu
contents Process Reward Models (PRMs) provide step-level supervision to large language models (LLMs), but scaling up training data annotation remains challenging for both humans and LLMs. To address this limitation, we propose an active learning approach, ActPRM, which proactively selects the most uncertain samples for training, substantially reducing labeling costs. During training, we use the PRM to estimate uncertainty after the forward pass, retaining only highly uncertain data. A capable yet costly reasoning model then labels this data. Then we compute the loss with respect to the labels and update the PRM's weights. We compare ActPRM vs. vanilla fine-tuning, on a pool-based active learning setting, demonstrating that ActPRM reduces 50% annotation, but achieving the comparable or even better performance. Beyond annotation efficiency, we further advance the actively trained PRM by filtering over 1M+ math reasoning trajectories with ActPRM, retaining 60% of the data. A subsequent training on this selected dataset yields a new state-of-the-art (SOTA) PRM on ProcessBench (75.0%) and PRMBench (65.5%) compared with same sized models.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10559
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Process Reward Model Training via Active Learning
Duan, Keyu
Liu, Zichen
Mao, Xin
Pang, Tianyu
Chen, Changyu
Chen, Qiguang
Shieh, Michael Qizhe
Dou, Longxu
Machine Learning
Artificial Intelligence
Process Reward Models (PRMs) provide step-level supervision to large language models (LLMs), but scaling up training data annotation remains challenging for both humans and LLMs. To address this limitation, we propose an active learning approach, ActPRM, which proactively selects the most uncertain samples for training, substantially reducing labeling costs. During training, we use the PRM to estimate uncertainty after the forward pass, retaining only highly uncertain data. A capable yet costly reasoning model then labels this data. Then we compute the loss with respect to the labels and update the PRM's weights. We compare ActPRM vs. vanilla fine-tuning, on a pool-based active learning setting, demonstrating that ActPRM reduces 50% annotation, but achieving the comparable or even better performance. Beyond annotation efficiency, we further advance the actively trained PRM by filtering over 1M+ math reasoning trajectories with ActPRM, retaining 60% of the data. A subsequent training on this selected dataset yields a new state-of-the-art (SOTA) PRM on ProcessBench (75.0%) and PRMBench (65.5%) compared with same sized models.
title Efficient Process Reward Model Training via Active Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2504.10559