DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chan, Chi-Min, Hajiramezanali, Ehsan, Li, Xiner, De Brouwer, Edward, Edwards, Carl, Xue, Wei, Han, Sirui, Guo, Yike, Scalia, Gabriele
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917324366282752
author Chan, Chi-Min
Hajiramezanali, Ehsan
Li, Xiner
De Brouwer, Edward
Edwards, Carl
Xue, Wei
Han, Sirui
Guo, Yike
Scalia, Gabriele
author_facet Chan, Chi-Min
Hajiramezanali, Ehsan
Li, Xiner
De Brouwer, Edward
Edwards, Carl
Xue, Wei
Han, Sirui
Guo, Yike
Scalia, Gabriele
contents In scientific reasoning tasks, the veracity of the reasoning process is as critical as the final outcome. While Process Reward Models (PRMs) offer a solution to the coarse-grained supervision problems inherent in Outcome Reward Models (ORMs), their deployment is hindered by the prohibitive cost of obtaining expert-verified step-wise labels. This paper addresses the challenge of training reliable PRMs using abundant but noisy "weak" supervision. We argue that existing Weak-to-Strong Generalization (W2SG) theories lack prescriptive guidelines for selecting high-quality training signals from noisy data. To bridge this gap, we introduce the Dual-Consensus Weak-to-Strong (DC-W2S) framework. By intersecting Self-Consensus (SC) metrics among weak supervisors with Neighborhood-Consensus (NC) metrics in the embedding space, we stratify supervision signals into distinct reliability regimes. We then employ a curriculum of instance-level balanced sampling and label-level reliability-aware masking to guide the training process. We demonstrate that DC-W2S enables the training of robust PRMs for complex reasoning without exhaustive expert annotation, proving that strategic data curation is more effective than indiscriminate training on large-scale noisy datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08095
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
Chan, Chi-Min
Hajiramezanali, Ehsan
Li, Xiner
De Brouwer, Edward
Edwards, Carl
Xue, Wei
Han, Sirui
Guo, Yike
Scalia, Gabriele
Computation and Language
Artificial Intelligence
Machine Learning
In scientific reasoning tasks, the veracity of the reasoning process is as critical as the final outcome. While Process Reward Models (PRMs) offer a solution to the coarse-grained supervision problems inherent in Outcome Reward Models (ORMs), their deployment is hindered by the prohibitive cost of obtaining expert-verified step-wise labels. This paper addresses the challenge of training reliable PRMs using abundant but noisy "weak" supervision. We argue that existing Weak-to-Strong Generalization (W2SG) theories lack prescriptive guidelines for selecting high-quality training signals from noisy data. To bridge this gap, we introduce the Dual-Consensus Weak-to-Strong (DC-W2S) framework. By intersecting Self-Consensus (SC) metrics among weak supervisors with Neighborhood-Consensus (NC) metrics in the embedding space, we stratify supervision signals into distinct reliability regimes. We then employ a curriculum of instance-level balanced sampling and label-level reliability-aware masking to guide the training process. We demonstrate that DC-W2S enables the training of robust PRMs for complex reasoning without exhaustive expert annotation, proving that strategic data curation is more effective than indiscriminate training on large-scale noisy datasets.
title DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.08095