Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yifei, Chakraborty, Tusher, Sharma, Srinagesh, Nunes, Leonardo, Sharma, Swati, Demopulos, Kate Drakos, Kıcıman, Emre, Lu, Songwu, Chandra, Ranveer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911659550834688
author Xu, Yifei
Chakraborty, Tusher
Sharma, Srinagesh
Nunes, Leonardo
Sharma, Swati
Demopulos, Kate Drakos
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
author_facet Xu, Yifei
Chakraborty, Tusher
Sharma, Srinagesh
Nunes, Leonardo
Sharma, Swati
Demopulos, Kate Drakos
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
contents Reinforcement learning (RL) training of large language models (LLMs) on unverifiable tasks is challenging even when a reasonable-quality reference answer is available. We propose a constrained RL training framework that (i) optimizes a token-level dense Reasoning Reflection Reward (R3) aligned with reasoning quality, and (ii) enforces rubric-gating as feasibility constraints at the rollout group level. R3 measures the model's token-level certainty of a reference answer under its chain-of-thought (CoT) prefix, and selectively emphasizes tokens with high cross-rollout variance, which we call reasoning-reflective tokens, that would otherwise be diluted by the bulk of low-variance tokens. The same variance signal also drives a filter that discards queries with insufficient signal for comparative learning. Rubric-gating complements R3 by operationalizing principled task criteria as hard accept/reject checks on final answers. Empirically, across four datasets spanning scientific writing, medicine, legal contracts, and finance, our framework outperforms strong baselines, achieves faster, more sample-efficient learning, and respects feasibility constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13351
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
Xu, Yifei
Chakraborty, Tusher
Sharma, Srinagesh
Nunes, Leonardo
Sharma, Swati
Demopulos, Kate Drakos
Kıcıman, Emre
Lu, Songwu
Chandra, Ranveer
Computation and Language
Artificial Intelligence
Machine Learning
Reinforcement learning (RL) training of large language models (LLMs) on unverifiable tasks is challenging even when a reasonable-quality reference answer is available. We propose a constrained RL training framework that (i) optimizes a token-level dense Reasoning Reflection Reward (R3) aligned with reasoning quality, and (ii) enforces rubric-gating as feasibility constraints at the rollout group level. R3 measures the model's token-level certainty of a reference answer under its chain-of-thought (CoT) prefix, and selectively emphasizes tokens with high cross-rollout variance, which we call reasoning-reflective tokens, that would otherwise be diluted by the bulk of low-variance tokens. The same variance signal also drives a filter that discards queries with insufficient signal for comparative learning. Rubric-gating complements R3 by operationalizing principled task criteria as hard accept/reject checks on final answers. Empirically, across four datasets spanning scientific writing, medicine, legal contracts, and finance, our framework outperforms strong baselines, achieves faster, more sample-efficient learning, and respects feasibility constraints.
title Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.13351