RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Daniel, Stante, Samuel, Redhardt, Florian, Libon, Lena, Kassraie, Parnian, Hakimi, Ido, Pásztor, Barna, Krause, Andreas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918360023826432
author Yang, Daniel
Stante, Samuel
Redhardt, Florian
Libon, Lena
Kassraie, Parnian
Hakimi, Ido
Pásztor, Barna
Krause, Andreas
author_facet Yang, Daniel
Stante, Samuel
Redhardt, Florian
Libon, Lena
Kassraie, Parnian
Hakimi, Ido
Pásztor, Barna
Krause, Andreas
contents Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncertainty in reward models arising from limited human feedback. Recent work suggests that quantifying this uncertainty can reduce the costs of human annotation via uncertainty-guided active learning and mitigate reward overoptimization in LLM post-training. However, uncertainty-aware reward models have so far been adopted without thorough comparison, leaving them poorly understood. This work introduces a unified framework, RewardUQ, to systematically evaluate uncertainty quantification for reward models. We compare common methods along standard metrics measuring accuracy and calibration, and we propose a new ranking strategy incorporating both dimensions for a simplified comparison. Our experimental results suggest that model size and initialization have the most meaningful impact on performance, and most prior work could have benefited from alternative design choices. To foster the development and evaluation of new methods and aid the deployment in downstream applications, we release our open-source framework as a Python package. Our code is available at https://github.com/lasgroup/rewarduq.
format Preprint
id arxiv_https___arxiv_org_abs_2602_24040
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
Yang, Daniel
Stante, Samuel
Redhardt, Florian
Libon, Lena
Kassraie, Parnian
Hakimi, Ido
Pásztor, Barna
Krause, Andreas
Machine Learning
Artificial Intelligence
Computation and Language
Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncertainty in reward models arising from limited human feedback. Recent work suggests that quantifying this uncertainty can reduce the costs of human annotation via uncertainty-guided active learning and mitigate reward overoptimization in LLM post-training. However, uncertainty-aware reward models have so far been adopted without thorough comparison, leaving them poorly understood. This work introduces a unified framework, RewardUQ, to systematically evaluate uncertainty quantification for reward models. We compare common methods along standard metrics measuring accuracy and calibration, and we propose a new ranking strategy incorporating both dimensions for a simplified comparison. Our experimental results suggest that model size and initialization have the most meaningful impact on performance, and most prior work could have benefited from alternative design choices. To foster the development and evaluation of new methods and aid the deployment in downstream applications, we release our open-source framework as a Python package. Our code is available at https://github.com/lasgroup/rewarduq.
title RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2602.24040