Reward Reasoning Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Jiaxin, Chi, Zewen, Dong, Li, Dong, Qingxiu, Wu, Xun, Huang, Shaohan, Wei, Furu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912384390529024
author Guo, Jiaxin
Chi, Zewen
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
author_facet Guo, Jiaxin
Chi, Zewen
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
contents Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In this work, we introduce Reward Reasoning Models (RRMs), which are specifically designed to execute a deliberate reasoning process before generating final rewards. Through chain-of-thought reasoning, RRMs leverage additional test-time compute for complex queries where appropriate rewards are not immediately apparent. To develop RRMs, we implement a reinforcement learning framework that fosters self-evolved reward reasoning capabilities without requiring explicit reasoning traces as training data. Experimental results demonstrate that RRMs achieve superior performance on reward modeling benchmarks across diverse domains. Notably, we show that RRMs can adaptively exploit test-time compute to further improve reward accuracy. The pretrained reward reasoning models are available at https://huggingface.co/Reward-Reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reward Reasoning Model
Guo, Jiaxin
Chi, Zewen
Dong, Li
Dong, Qingxiu
Wu, Xun
Huang, Shaohan
Wei, Furu
Computation and Language
Reward models play a critical role in guiding large language models toward outputs that align with human expectations. However, an open challenge remains in effectively utilizing test-time compute to enhance reward model performance. In this work, we introduce Reward Reasoning Models (RRMs), which are specifically designed to execute a deliberate reasoning process before generating final rewards. Through chain-of-thought reasoning, RRMs leverage additional test-time compute for complex queries where appropriate rewards are not immediately apparent. To develop RRMs, we implement a reinforcement learning framework that fosters self-evolved reward reasoning capabilities without requiring explicit reasoning traces as training data. Experimental results demonstrate that RRMs achieve superior performance on reward modeling benchmarks across diverse domains. Notably, we show that RRMs can adaptively exploit test-time compute to further improve reward accuracy. The pretrained reward reasoning models are available at https://huggingface.co/Reward-Reasoning.
title Reward Reasoning Model
topic Computation and Language
url https://arxiv.org/abs/2505.14674