Reinforcement Learning with Rubric Anchors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zenan, Zhuang, Yihong, Lu, Guoshan, Qin, Zeyu, Xu, Haokai, Zhao, Tianyu, Peng, Ru, Hu, Jiaqi, Shen, Zhanming, Hu, Xiaomeng, Gu, Xijun, Tu, Peiyi, Liu, Jiaxin, Chen, Wenyu, Fu, Yuzhuo, Fan, Zhiting, Gu, Yanmei, Wang, Yuanyuan, Yang, Zhengkai, Li, Jianguo, Zhao, Junbo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915449243959296
author Huang, Zenan
Zhuang, Yihong
Lu, Guoshan
Qin, Zeyu
Xu, Haokai
Zhao, Tianyu
Peng, Ru
Hu, Jiaqi
Shen, Zhanming
Hu, Xiaomeng
Gu, Xijun
Tu, Peiyi
Liu, Jiaxin
Chen, Wenyu
Fu, Yuzhuo
Fan, Zhiting
Gu, Yanmei
Wang, Yuanyuan
Yang, Zhengkai
Li, Jianguo
Zhao, Junbo
author_facet Huang, Zenan
Zhuang, Yihong
Lu, Guoshan
Qin, Zeyu
Xu, Haokai
Zhao, Tianyu
Peng, Ru
Hu, Jiaqi
Shen, Zhanming
Hu, Xiaomeng
Gu, Xijun
Tu, Peiyi
Liu, Jiaxin
Chen, Wenyu
Fu, Yuzhuo
Fan, Zhiting
Gu, Yanmei
Wang, Yuanyuan
Yang, Zhengkai
Li, Jianguo
Zhao, Junbo
contents Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing Large Language Models (LLMs), exemplified by the success of OpenAI's o-series. In RLVR, rewards are derived from verifiable signals-such as passing unit tests in code generation or matching correct answers in mathematical reasoning. While effective, this requirement largely confines RLVR to domains with automatically checkable outcomes. To overcome this, we extend the RLVR paradigm to open-ended tasks by integrating rubric-based rewards, where carefully designed rubrics serve as structured, model-interpretable criteria for automatic scoring of subjective outputs. We construct, to our knowledge, the largest rubric reward system to date, with over 10,000 rubrics from humans, LLMs, or a hybrid human-LLM collaboration. Implementing rubric-based RL is challenging; we tackle these issues with a clear framework and present an open-sourced Qwen-30B-A3B model with notable gains: 1) With only 5K+ samples, our system improves by +5.2% on open-ended benchmarks (especially humanities), outperforming a 671B DeepSeek-V3 model by +2.4%, while preserving general and reasoning abilities. 2) Our method provides fine-grained stylistic control, using rubrics as anchors to mitigate the "AI-like" tone and produce more human-like, expressive responses. We share key lessons in rubric construction, data selection, and training, and discuss limitations and future releases.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12790
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning with Rubric Anchors
Huang, Zenan
Zhuang, Yihong
Lu, Guoshan
Qin, Zeyu
Xu, Haokai
Zhao, Tianyu
Peng, Ru
Hu, Jiaqi
Shen, Zhanming
Hu, Xiaomeng
Gu, Xijun
Tu, Peiyi
Liu, Jiaxin
Chen, Wenyu
Fu, Yuzhuo
Fan, Zhiting
Gu, Yanmei
Wang, Yuanyuan
Yang, Zhengkai
Li, Jianguo
Zhao, Junbo
Artificial Intelligence
Computation and Language
Machine Learning
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing Large Language Models (LLMs), exemplified by the success of OpenAI's o-series. In RLVR, rewards are derived from verifiable signals-such as passing unit tests in code generation or matching correct answers in mathematical reasoning. While effective, this requirement largely confines RLVR to domains with automatically checkable outcomes. To overcome this, we extend the RLVR paradigm to open-ended tasks by integrating rubric-based rewards, where carefully designed rubrics serve as structured, model-interpretable criteria for automatic scoring of subjective outputs. We construct, to our knowledge, the largest rubric reward system to date, with over 10,000 rubrics from humans, LLMs, or a hybrid human-LLM collaboration. Implementing rubric-based RL is challenging; we tackle these issues with a clear framework and present an open-sourced Qwen-30B-A3B model with notable gains: 1) With only 5K+ samples, our system improves by +5.2% on open-ended benchmarks (especially humanities), outperforming a 671B DeepSeek-V3 model by +2.4%, while preserving general and reasoning abilities. 2) Our method provides fine-grained stylistic control, using rubrics as anchors to mitigate the "AI-like" tone and produce more human-like, expressive responses. We share key lessons in rubric construction, data selection, and training, and discuss limitations and future releases.
title Reinforcement Learning with Rubric Anchors
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2508.12790