Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Lipeng, Huang, Sen, Zhang, Zhuo, Zou, Anni, Zhai, Yunpeng, Ren, Dingchao, Zhang, Kezun, Hu, Haoyuan, Liu, Boyin, Chen, Haoran, Liu, Zhaoyang, Ding, Bolin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911423548882944
author Xie, Lipeng
Huang, Sen
Zhang, Zhuo
Zou, Anni
Zhai, Yunpeng
Ren, Dingchao
Zhang, Kezun
Hu, Haoyuan
Liu, Boyin
Chen, Haoran
Liu, Zhaoyang
Ding, Bolin
author_facet Xie, Lipeng
Huang, Sen
Zhang, Zhuo
Zou, Anni
Zhai, Yunpeng
Ren, Dingchao
Zhang, Kezun
Hu, Haoyuan
Liu, Boyin
Chen, Haoran
Liu, Zhaoyang
Ding, Bolin
contents Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit reward parameterization, recasting optimization from continuous weight spaces to the discrete space of natural language rubrics. We introduce a training-free framework based on iterative rubric learning: it locally induces discriminative criteria via verification-driven refinement, and globally compresses the candidate criteria pool into a compact core set by maximizing an information-theoretic coding rate objective. We organize the compressed core set into a hierarchical rubric structure -- high-level evaluation dimensions supported by concrete verification checks -- serving as an interpretable, portable reward function. Empirically, our approach challenges prevailing data scaling assumptions: using only 70 preference pairs, our rubric-guided judges outperform fully trained reward models on diverse benchmarks. For instance, Qwen3-8B equipped with our learned rubrics achieves 80.91% on RewardBench2, surpassing the specialized Skywork-Reward-V2-Qwen3-8B (78.20%). These results demonstrate that alignment signals are highly compressible and can be effectively captured through explicit symbolic search.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17314
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
Xie, Lipeng
Huang, Sen
Zhang, Zhuo
Zou, Anni
Zhai, Yunpeng
Ren, Dingchao
Zhang, Kezun
Hu, Haoyuan
Liu, Boyin
Chen, Haoran
Liu, Zhaoyang
Ding, Bolin
Machine Learning
Artificial Intelligence
Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit reward parameterization, recasting optimization from continuous weight spaces to the discrete space of natural language rubrics. We introduce a training-free framework based on iterative rubric learning: it locally induces discriminative criteria via verification-driven refinement, and globally compresses the candidate criteria pool into a compact core set by maximizing an information-theoretic coding rate objective. We organize the compressed core set into a hierarchical rubric structure -- high-level evaluation dimensions supported by concrete verification checks -- serving as an interpretable, portable reward function. Empirically, our approach challenges prevailing data scaling assumptions: using only 70 preference pairs, our rubric-guided judges outperform fully trained reward models on diverse benchmarks. For instance, Qwen3-8B equipped with our learned rubrics achieves 80.91% on RewardBench2, surpassing the specialized Skywork-Reward-V2-Qwen3-8B (78.20%). These results demonstrate that alignment signals are highly compressible and can be effectively captured through explicit symbolic search.
title Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2510.17314