Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lou, Xingzhou, Yan, Dong, Shen, Wei, Yan, Yuzi, Xie, Jian, Zhang, Junge
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909488837033984
author Lou, Xingzhou
Yan, Dong
Shen, Wei
Yan, Yuzi
Xie, Jian
Zhang, Junge
author_facet Lou, Xingzhou
Yan, Dong
Shen, Wei
Yan, Yuzi
Xie, Jian
Zhang, Junge
contents Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of human preferences and fail to assess the reliability of reward predictions. To address these challenges, we introduce the Uncertainty-aware Reward Model (URM) and its ensemble variant, URME. URM employs a probabilistic value head to capture aleatoric uncertainty by modeling the distribution of disentangled human preference attributes. URME further quantifies epistemic uncertainty by examining discrepancies among individual URMs within the ensemble, enabling identification of unreliable evaluations. Our empirical evaluations demonstrate that URM achieves strong performance on RewardBench, outperforming competitive large-scale models. Additionally, extensive experiments, including best-of-n sampling (BoN), iterative direct preference optimization (iterative DPO), and proximal policy optimization (PPO), demonstrate that URM and URME significantly enhance LLMs' generation quality. Notably, reward predictions with lower uncertainty are far more reliable, demonstrate significantly higher quality, and result in substantially improved alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00847
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
Lou, Xingzhou
Yan, Dong
Shen, Wei
Yan, Yuzi
Xie, Jian
Zhang, Junge
Machine Learning
Reward models (RMs) are essential for aligning large language models (LLM) with human expectations. However, existing RMs struggle to capture the stochastic and uncertain nature of human preferences and fail to assess the reliability of reward predictions. To address these challenges, we introduce the Uncertainty-aware Reward Model (URM) and its ensemble variant, URME. URM employs a probabilistic value head to capture aleatoric uncertainty by modeling the distribution of disentangled human preference attributes. URME further quantifies epistemic uncertainty by examining discrepancies among individual URMs within the ensemble, enabling identification of unreliable evaluations. Our empirical evaluations demonstrate that URM achieves strong performance on RewardBench, outperforming competitive large-scale models. Additionally, extensive experiments, including best-of-n sampling (BoN), iterative direct preference optimization (iterative DPO), and proximal policy optimization (PPO), demonstrate that URM and URME significantly enhance LLMs' generation quality. Notably, reward predictions with lower uncertainty are far more reliable, demonstrate significantly higher quality, and result in substantially improved alignment.
title Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown
topic Machine Learning
url https://arxiv.org/abs/2410.00847