Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bhattarai, Manish, Boureima, Ismael, Ranasinghe, Nishath Rajiv, Pakin, Scott, O'Malley, Dan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918489929809920
author Bhattarai, Manish
Boureima, Ismael
Ranasinghe, Nishath Rajiv
Pakin, Scott
O'Malley, Dan
author_facet Bhattarai, Manish
Boureima, Ismael
Ranasinghe, Nishath Rajiv
Pakin, Scott
O'Malley, Dan
contents We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal: instead of a binary outcome or a single holistic score, each response is graded along multiple task-specific criteria. We formalize \emph{rubric-grounded reinforcement learning (RL)}: a framework in which the policy is optimized against a structured, multi-criterion reward produced by a frozen LLM judge that conditions on auxiliary grounding the policy never sees. We instantiate the framework by deriving rubrics from an Office of Scientific and Technical Information (OSTI)-derived corpus of roughly 100,000 scientific and technical documents and training Llama-3.1-8B-Instruct with Group Relative Policy Optimization (GRPO). With GRPO-based training, the model achieves $71.7\%$ normalized reward on held-out rubric evaluation. The GRPO-tuned policy also improves over the base model on four reasoning benchmarks not derived from the training corpus -- GSM8K, MATH, GPQA Main, and GPQA Diamond. These results provide evidence that structured, document-grounded rewards can improve held-out rubric performance and induce transferable reasoning behaviors beyond the corpus used to construct the training environment.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08061
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
Bhattarai, Manish
Boureima, Ismael
Ranasinghe, Nishath Rajiv
Pakin, Scott
O'Malley, Dan
Artificial Intelligence
We argue that decomposing reward into weighted, verifiable criteria and using an LLM judge to score them provides a partial-credit optimization signal: instead of a binary outcome or a single holistic score, each response is graded along multiple task-specific criteria. We formalize \emph{rubric-grounded reinforcement learning (RL)}: a framework in which the policy is optimized against a structured, multi-criterion reward produced by a frozen LLM judge that conditions on auxiliary grounding the policy never sees. We instantiate the framework by deriving rubrics from an Office of Scientific and Technical Information (OSTI)-derived corpus of roughly 100,000 scientific and technical documents and training Llama-3.1-8B-Instruct with Group Relative Policy Optimization (GRPO). With GRPO-based training, the model achieves $71.7\%$ normalized reward on held-out rubric evaluation. The GRPO-tuned policy also improves over the base model on four reasoning benchmarks not derived from the training corpus -- GSM8K, MATH, GPQA Main, and GPQA Diamond. These results provide evidence that structured, document-grounded rewards can improve held-out rubric performance and induce transferable reasoning behaviors beyond the corpus used to construct the training environment.
title Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2605.08061