Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911423548882944 |
|---|---|
| author | Xie, Lipeng Huang, Sen Zhang, Zhuo Zou, Anni Zhai, Yunpeng Ren, Dingchao Zhang, Kezun Hu, Haoyuan Liu, Boyin Chen, Haoran Liu, Zhaoyang Ding, Bolin |
| author_facet | Xie, Lipeng Huang, Sen Zhang, Zhuo Zou, Anni Zhai, Yunpeng Ren, Dingchao Zhang, Kezun Hu, Haoyuan Liu, Boyin Chen, Haoran Liu, Zhaoyang Ding, Bolin |
| contents | Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit reward parameterization, recasting optimization from continuous weight spaces to the discrete space of natural language rubrics. We introduce a training-free framework based on iterative rubric learning: it locally induces discriminative criteria via verification-driven refinement, and globally compresses the candidate criteria pool into a compact core set by maximizing an information-theoretic coding rate objective. We organize the compressed core set into a hierarchical rubric structure -- high-level evaluation dimensions supported by concrete verification checks -- serving as an interpretable, portable reward function. Empirically, our approach challenges prevailing data scaling assumptions: using only 70 preference pairs, our rubric-guided judges outperform fully trained reward models on diverse benchmarks. For instance, Qwen3-8B equipped with our learned rubrics achieves 80.91% on RewardBench2, surpassing the specialized Skywork-Reward-V2-Qwen3-8B (78.20%). These results demonstrate that alignment signals are highly compressible and can be effectively captured through explicit symbolic search. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_17314 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling Xie, Lipeng Huang, Sen Zhang, Zhuo Zou, Anni Zhai, Yunpeng Ren, Dingchao Zhang, Kezun Hu, Haoyuan Liu, Boyin Chen, Haoran Liu, Zhaoyang Ding, Bolin Machine Learning Artificial Intelligence Conventional reward modeling relies on gradient descent over neural weights, creating opaque, data-hungry "black boxes." We propose a paradigm shift from implicit to explicit reward parameterization, recasting optimization from continuous weight spaces to the discrete space of natural language rubrics. We introduce a training-free framework based on iterative rubric learning: it locally induces discriminative criteria via verification-driven refinement, and globally compresses the candidate criteria pool into a compact core set by maximizing an information-theoretic coding rate objective. We organize the compressed core set into a hierarchical rubric structure -- high-level evaluation dimensions supported by concrete verification checks -- serving as an interpretable, portable reward function. Empirically, our approach challenges prevailing data scaling assumptions: using only 70 preference pairs, our rubric-guided judges outperform fully trained reward models on diverse benchmarks. For instance, Qwen3-8B equipped with our learned rubrics achieves 80.91% on RewardBench2, surpassing the specialized Skywork-Reward-V2-Qwen3-8B (78.20%). These results demonstrate that alignment signals are highly compressible and can be effectively captured through explicit symbolic search. |
| title | Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2510.17314 |