RM-R1: Reward Modeling as Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Xiusi, Li, Gaotang, Wang, Ziqi, Jin, Bowen, Qian, Cheng, Wang, Yu, Wang, Hongru, Zhang, Yu, Zhang, Denghui, Zhang, Tong, Tong, Hanghang, Ji, Heng
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908869361401856
author Chen, Xiusi
Li, Gaotang
Wang, Ziqi
Jin, Bowen
Qian, Cheng
Wang, Yu
Wang, Hongru
Zhang, Yu
Zhang, Denghui
Zhang, Tong
Tong, Hanghang
Ji, Heng
author_facet Chen, Xiusi
Li, Gaotang
Wang, Ziqi
Jin, Bowen
Qian, Cheng
Wang, Yu
Wang, Hongru
Zhang, Yu
Zhang, Denghui
Zhang, Tong
Tong, Hanghang
Ji, Heng
contents Reward modeling is essential for aligning large language models with human preferences through reinforcement learning. To provide accurate reward signals, a reward model (RM) should stimulate deep thinking and conduct interpretable reasoning before assigning a score or a judgment. Inspired by recent advances of long chain-of-thought on reasoning-intensive tasks, we hypothesize and validate that integrating reasoning into reward modeling significantly enhances RM's interpretability and performance. We introduce a new class of generative reward models, Reasoning Reward Models (ReasRMs), which formulate reward modeling as a reasoning task. We propose a reasoning-oriented training pipeline and train a family of ReasRMs, RM-R1. RM-R1 features a chain-of-rubrics (CoR) mechanism -- self-generating sample-level chat rubrics or math/code solutions, and evaluating candidate responses against them. The training of RM-R1 consists of two key stages: (1) distillation of high-quality reasoning chains and (2) reinforcement learning with verifiable rewards. Empirically, our models achieve superior performance across three reward model benchmarks on average, outperforming much larger open-weight models (e.g., INF-ORM-Llama3.1-70B) and proprietary ones (e.g., GPT-4o) by up to 4.9%. Beyond final performance, we perform thorough analyses to understand the key ingredients of successful ReasRM training.
format Preprint
id arxiv_https___arxiv_org_abs_2505_02387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RM-R1: Reward Modeling as Reasoning
Chen, Xiusi
Li, Gaotang
Wang, Ziqi
Jin, Bowen
Qian, Cheng
Wang, Yu
Wang, Hongru
Zhang, Yu
Zhang, Denghui
Zhang, Tong
Tong, Hanghang
Ji, Heng
Computation and Language
Artificial Intelligence
Machine Learning
Reward modeling is essential for aligning large language models with human preferences through reinforcement learning. To provide accurate reward signals, a reward model (RM) should stimulate deep thinking and conduct interpretable reasoning before assigning a score or a judgment. Inspired by recent advances of long chain-of-thought on reasoning-intensive tasks, we hypothesize and validate that integrating reasoning into reward modeling significantly enhances RM's interpretability and performance. We introduce a new class of generative reward models, Reasoning Reward Models (ReasRMs), which formulate reward modeling as a reasoning task. We propose a reasoning-oriented training pipeline and train a family of ReasRMs, RM-R1. RM-R1 features a chain-of-rubrics (CoR) mechanism -- self-generating sample-level chat rubrics or math/code solutions, and evaluating candidate responses against them. The training of RM-R1 consists of two key stages: (1) distillation of high-quality reasoning chains and (2) reinforcement learning with verifiable rewards. Empirically, our models achieve superior performance across three reward model benchmarks on average, outperforming much larger open-weight models (e.g., INF-ORM-Llama3.1-70B) and proprietary ones (e.g., GPT-4o) by up to 4.9%. Beyond final performance, we perform thorough analyses to understand the key ingredients of successful ReasRM training.
title RM-R1: Reward Modeling as Reasoning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.02387