CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Dengcan, Yang, Fengkai, Wang, Xiaohan, Yan, Shurui, Chai, Jiajun, Li, Jiahao, Ban, Yikun, Mao, Zhendong, Lin, Wei, Yin, Guojun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910046170906624
author Liu, Dengcan
Yang, Fengkai
Wang, Xiaohan
Yan, Shurui
Chai, Jiajun
Li, Jiahao
Ban, Yikun
Mao, Zhendong
Lin, Wei
Yin, Guojun
author_facet Liu, Dengcan
Yang, Fengkai
Wang, Xiaohan
Yan, Shurui
Chai, Jiajun
Li, Jiahao
Ban, Yikun
Mao, Zhendong
Lin, Wei
Yin, Guojun
contents Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on costly expert annotations. While recent rubric-based approaches enhance evaluation transparency, they lack systematic quality control, yielding noisy and redundant criteria, failing to mitigate persistent biases (e.g., verbosity, position) in LLM evaluators, and creating a scalability-reliability trade-off. To address these limitations, we propose CDRRM (Contrast-Driven Rubric Reward Model), a framework built on a novel Contrast-then-Synthesis paradigm for high-quality rubric generation and guided preference judgment. CDRRM first conducts multi-dimensional contrastive profiling on preference pairs to identify causal discriminative factors, then synthesizes these insights into compact, context-aware rubrics to guide preference judg- ments. Extensive experiments on three authoritative benchmarks (RewardBench, RMBench, RMB) demonstrate that CDRRM achieves state-of-the-art performance across diverse domains and effectively mitigates aforementioned evaluation biases. Notably, our approach delivers exceptional data efficiency: training the rubric generator on only 3k high-quality samples empowers a frozen pre-trained judge model to outperform fully fine-tuned baselines. This work offers a scalable, interpretable, and data-efficient path for reward modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08035
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
Liu, Dengcan
Yang, Fengkai
Wang, Xiaohan
Yan, Shurui
Chai, Jiajun
Li, Jiahao
Ban, Yikun
Mao, Zhendong
Lin, Wei
Yin, Guojun
Artificial Intelligence
Machine Learning
Reward modeling is essential for aligning Large Language Models(LLMs) with human preferences, yet conventional reward models suffer from poor interpretability and heavy reliance on costly expert annotations. While recent rubric-based approaches enhance evaluation transparency, they lack systematic quality control, yielding noisy and redundant criteria, failing to mitigate persistent biases (e.g., verbosity, position) in LLM evaluators, and creating a scalability-reliability trade-off. To address these limitations, we propose CDRRM (Contrast-Driven Rubric Reward Model), a framework built on a novel Contrast-then-Synthesis paradigm for high-quality rubric generation and guided preference judgment. CDRRM first conducts multi-dimensional contrastive profiling on preference pairs to identify causal discriminative factors, then synthesizes these insights into compact, context-aware rubrics to guide preference judg- ments. Extensive experiments on three authoritative benchmarks (RewardBench, RMBench, RMB) demonstrate that CDRRM achieves state-of-the-art performance across diverse domains and effectively mitigates aforementioned evaluation biases. Notably, our approach delivers exceptional data efficiency: training the rubric generator on only 3k high-quality samples empowers a frozen pre-trained judge model to outperform fully fine-tuned baselines. This work offers a scalable, interpretable, and data-efficient path for reward modeling.
title CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.08035