GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Chenglong, Mu, Yongyu, Zhou, Hang, Huo, Yifu, Zhu, Ziming, Zeng, Jiali, Yang, Murun, Li, Bei, Hao, Xiaoyang, Zhang, Chunliang, Meng, Fandong, Zhu, Jingbo, Xiao, Tong
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912711836696576
author Wang, Chenglong
Mu, Yongyu
Zhou, Hang
Huo, Yifu
Zhu, Ziming
Zeng, Jiali
Yang, Murun
Li, Bei
Hao, Xiaoyang
Zhang, Chunliang
Meng, Fandong
Zhu, Jingbo
Xiao, Tong
author_facet Wang, Chenglong
Mu, Yongyu
Zhou, Hang
Huo, Yifu
Zhu, Ziming
Zeng, Jiali
Yang, Murun
Li, Bei
Hao, Xiaoyang
Zhang, Chunliang
Meng, Fandong
Zhu, Jingbo
Xiao, Tong
contents Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short of instilling explicit reasoning into reward models. To bridge this gap, we propose a self-training approach that leverages unlabeled data to elicit reward reasoning in reward models. Based on this approach, we develop GRAM-R$^2$, a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R$^2$ can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as response ranking and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R$^2$ consistently delivers strong performance, outperforming several strong discriminative and generative baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_02492
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
Wang, Chenglong
Mu, Yongyu
Zhou, Hang
Huo, Yifu
Zhu, Ziming
Zeng, Jiali
Yang, Murun
Li, Bei
Hao, Xiaoyang
Zhang, Chunliang
Meng, Fandong
Zhu, Jingbo
Xiao, Tong
Computation and Language
Machine Learning
Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, developing effective reward models remains a fundamental challenge: the heavy reliance on large-scale labeled preference data. Pre-training on abundant unlabeled data offers a promising direction, but existing approaches fall short of instilling explicit reasoning into reward models. To bridge this gap, we propose a self-training approach that leverages unlabeled data to elicit reward reasoning in reward models. Based on this approach, we develop GRAM-R$^2$, a generative reward model trained to produce not only preference labels but also accompanying reward rationales. GRAM-R$^2$ can serve as a foundation model for reward reasoning and can be applied to a wide range of tasks with minimal or no additional fine-tuning. It can support downstream applications such as response ranking and task-specific reward tuning. Experiments on response ranking, task adaptation, and reinforcement learning from human feedback demonstrate that GRAM-R$^2$ consistently delivers strong performance, outperforming several strong discriminative and generative baselines.
title GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.02492