Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gunjal, Anisha, Wang, Anthony, Lau, Elaine, Nath, Vaskar, He, Yunzhong, Liu, Bing, Hendryx, Sean
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911190028910592
author Gunjal, Anisha
Wang, Anthony
Lau, Elaine
Nath, Vaskar
He, Yunzhong
Liu, Bing
Hendryx, Sean
author_facet Gunjal, Anisha
Wang, Anthony
Lau, Elaine
Nath, Vaskar
He, Yunzhong
Liu, Bing
Hendryx, Sean
contents Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than binary correctness. Instance-specific rubrics have recently been used in evaluation benchmarks to capture such judgments, but their potential as reward signals for on-policy post-training remains underexplored. We introduce $\textbf{Rubrics as Rewards}$ (RaR), an on-policy reinforcement learning method that extends RLVR beyond verifiable domains by using rubric-based feedback. Across both medical and science domains, we evaluate multiple strategies for aggregating rubric feedback into rewards. The best RaR variant achieves relative improvements of up to $31\%$ on HealthBench and $7\%$ on GPQA-Diamond over popular LLM-as-judge baselines that rely on direct Likert-based rewards. These results demonstrate that RaR-trained policies adapt well to diverse evaluation formats, performing strongly on both rubric-based and multiple-choice tasks. Moreover, we find that using rubrics as structured reward signals yields better alignment for smaller judges and reduces performance variance across judge scales.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17746
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Gunjal, Anisha
Wang, Anthony
Lau, Elaine
Nath, Vaskar
He, Yunzhong
Liu, Bing
Hendryx, Sean
Machine Learning
Artificial Intelligence
Computation and Language
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for complex reasoning tasks with clear correctness signals such as math and coding. However, extending it to real-world reasoning tasks is challenging, as evaluation depends on nuanced, multi-criteria judgments rather than binary correctness. Instance-specific rubrics have recently been used in evaluation benchmarks to capture such judgments, but their potential as reward signals for on-policy post-training remains underexplored. We introduce $\textbf{Rubrics as Rewards}$ (RaR), an on-policy reinforcement learning method that extends RLVR beyond verifiable domains by using rubric-based feedback. Across both medical and science domains, we evaluate multiple strategies for aggregating rubric feedback into rewards. The best RaR variant achieves relative improvements of up to $31\%$ on HealthBench and $7\%$ on GPQA-Diamond over popular LLM-as-judge baselines that rely on direct Likert-based rewards. These results demonstrate that RaR-trained policies adapt well to diverse evaluation formats, performing strongly on both rubric-based and multiple-choice tasks. Moreover, we find that using rubrics as structured reward signals yields better alignment for smaller judges and reduces performance variance across judge scales.
title Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.17746