Pre-Trained Policy Discriminators are General Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dou, Shihan, Liu, Shichun, Yang, Yuming, Zou, Yicheng, Zhou, Yunhua, Xing, Shuhao, Huang, Chenhao, Ge, Qiming, Song, Demin, Lv, Haijun, Gao, Songyang, Lv, Chengqi, Zhou, Enyu, Guo, Honglin, Xi, Zhiheng, Zhang, Wenwei, Guo, Qipeng, Zhang, Qi, Qiu, Xipeng, Huang, Xuanjing, Gui, Tao, Chen, Kai
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911385635520512
author Dou, Shihan
Liu, Shichun
Yang, Yuming
Zou, Yicheng
Zhou, Yunhua
Xing, Shuhao
Huang, Chenhao
Ge, Qiming
Song, Demin
Lv, Haijun
Gao, Songyang
Lv, Chengqi
Zhou, Enyu
Guo, Honglin
Xi, Zhiheng
Zhang, Wenwei
Guo, Qipeng
Zhang, Qi
Qiu, Xipeng
Huang, Xuanjing
Gui, Tao
Chen, Kai
author_facet Dou, Shihan
Liu, Shichun
Yang, Yuming
Zou, Yicheng
Zhou, Yunhua
Xing, Shuhao
Huang, Chenhao
Ge, Qiming
Song, Demin
Lv, Haijun
Gao, Songyang
Lv, Chengqi
Zhou, Enyu
Guo, Honglin
Xi, Zhiheng
Zhang, Wenwei
Guo, Qipeng
Zhang, Qi
Qiu, Xipeng
Huang, Xuanjing
Gui, Tao
Chen, Kai
contents We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named Policy Discriminative Learning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance--improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_05197
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pre-Trained Policy Discriminators are General Reward Models
Dou, Shihan
Liu, Shichun
Yang, Yuming
Zou, Yicheng
Zhou, Yunhua
Xing, Shuhao
Huang, Chenhao
Ge, Qiming
Song, Demin
Lv, Haijun
Gao, Songyang
Lv, Chengqi
Zhou, Enyu
Guo, Honglin
Xi, Zhiheng
Zhang, Wenwei
Guo, Qipeng
Zhang, Qi
Qiu, Xipeng
Huang, Xuanjing
Gui, Tao
Chen, Kai
Computation and Language
Machine Learning
We offer a novel perspective on reward modeling by formulating it as a policy discriminator, which quantifies the difference between two policies to generate a reward signal, guiding the training policy towards a target policy with desired behaviors. Based on this conceptual insight, we propose a scalable pre-training method named Policy Discriminative Learning (POLAR), which trains a reward model (RM) to discern identical policies and discriminate different ones. Unlike traditional reward modeling methods relying on absolute preferences, POLAR captures the relative difference between one policy and an arbitrary target policy, which is a scalable, high-level optimization objective suitable for modeling generic ranking relationships. Leveraging the POLAR pre-training paradigm, we present a series of RMs with parameter scales from 1.8B to 7B. Empirical results show that POLAR substantially outperforms traditional non-pre-trained methods, significantly enhancing RM performance. For instance, POLAR-7B could improve preference accuracy from 54.8% to 81.0% on STEM tasks and from 57.9% to 85.5% on creative writing tasks compared to SOTA baselines. POLAR also shows robust generalization capabilities in RLHF using Reinforcement Fine-tuning (RFT), providing reliable reward signals and markedly enhancing policy performance--improving LLaMa3.1-8B from an average of 47.36% to 56.33% and Qwen2.5-32B from 64.49% to 70.47% on 20 benchmarks. Moreover, scaling experiments reveal a clear power-law relationship between computation and performance, supported by linear correlation coefficients approaching 0.99. The impressive performance, strong generalization, and scaling properties suggest that POLAR is a promising direction for developing general and strong reward models.
title Pre-Trained Policy Discriminators are General Reward Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.05197