Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Rachel, Hadfield-Menell, Dylan, Greenewald, Kristjan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909037383122944
author Ma, Rachel
Hadfield-Menell, Dylan
Greenewald, Kristjan
author_facet Ma, Rachel
Hadfield-Menell, Dylan
Greenewald, Kristjan
contents Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the first use of conditional optimal transport for calibrating PRMs, modifying conditional OT (CondOT) map learning \cite{bunne2022supervised} to estimate a monotonic conditional quantile function over success probabilities estimated by the PRM, conditioned on PRM hidden states. This yields structurally valid quantile estimates and enables efficient extraction of confidence bounds at arbitrary levels, which we integrate into the instance-adaptive scaling (IAS) framework of \cite{park2025know}. We evaluate on mathematical reasoning benchmarks spanning moderate-difficulty problems (MATH-500) and harder out-of-distribution problems (AIME). For PRMs with reliable ranking signals, our method substantially improves calibration over both uncalibrated PRMs and quantile regression. On downstream Best-of-N IAS performance, our method generally improves over uncalibrated PRMs. These results establish conditional optimal transport as another principled and practical approach to PRM calibration, offering structural guarantees and flexible uncertainty estimation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06785
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
Ma, Rachel
Hadfield-Menell, Dylan
Greenewald, Kristjan
Machine Learning
Artificial Intelligence
Inference-time scaling methods rely on Process Reward Models (PRMs), which are often poorly calibrated and overestimate success probabilities. We propose, to our knowledge, the first use of conditional optimal transport for calibrating PRMs, modifying conditional OT (CondOT) map learning \cite{bunne2022supervised} to estimate a monotonic conditional quantile function over success probabilities estimated by the PRM, conditioned on PRM hidden states. This yields structurally valid quantile estimates and enables efficient extraction of confidence bounds at arbitrary levels, which we integrate into the instance-adaptive scaling (IAS) framework of \cite{park2025know}. We evaluate on mathematical reasoning benchmarks spanning moderate-difficulty problems (MATH-500) and harder out-of-distribution problems (AIME). For PRMs with reliable ranking signals, our method substantially improves calibration over both uncalibrated PRMs and quantile regression. On downstream Best-of-N IAS performance, our method generally improves over uncalibrated PRMs. These results establish conditional optimal transport as another principled and practical approach to PRM calibration, offering structural guarantees and flexible uncertainty estimation.
title Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.06785