RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Haoxiang, Dong, Zihan, Liu, Tianci, Wang, Wanying, Xu, Ran, Yu, Tony, Zhang, Linjun, Wang, Haoyu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914611909885952
author Jiang, Haoxiang
Dong, Zihan
Liu, Tianci
Wang, Wanying
Xu, Ran
Yu, Tony
Zhang, Linjun
Wang, Haoyu
author_facet Jiang, Haoxiang
Dong, Zihan
Liu, Tianci
Wang, Wanying
Xu, Ran
Yu, Tony
Zhang, Linjun
Wang, Haoyu
contents Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address this by decomposing evaluation into explicit criteria, but existing approaches typically depend on frontier LLMs and suffer from ties caused by hard Boolean aggregation. We present RUBRIC-ARROW, an alternating framework that jointly trains a rubric generator and a rubric-conditioned judge, with its RL stage using only pairwise preference data. Our method couples a probability-based scoring rule that reduces ties with phase-specific preference-based rewards and an alternating GRPO scheme that together train the pointwise evaluator. Extensive experiments show that RUBRIC-ARROW achieves competitive reward-modeling accuracy and yields consistent gains for downstream policy post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29156
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
Jiang, Haoxiang
Dong, Zihan
Liu, Tianci
Wang, Wanying
Xu, Ran
Yu, Tony
Zhang, Linjun
Wang, Haoyu
Machine Learning
Computation and Language
Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. Rubric-based methods address this by decomposing evaluation into explicit criteria, but existing approaches typically depend on frontier LLMs and suffer from ties caused by hard Boolean aggregation. We present RUBRIC-ARROW, an alternating framework that jointly trains a rubric generator and a rubric-conditioned judge, with its RL stage using only pairwise preference data. Our method couples a probability-based scoring rule that reduces ties with phase-specific preference-based rewards and an alternating GRPO scheme that together train the pointwise evaluator. Extensive experiments show that RUBRIC-ARROW achieves competitive reward-modeling accuracy and yields consistent gains for downstream policy post-training.
title RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2605.29156