Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lan, Guangchen, Xiong, Lian, Zhou, Xin, Cui, Hejie, Zhang, Yuwei, Li, Mao, Shi, Zhenyu, Fetahu, Besnik, Li, Lihong, Li, Xian
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918486626795520
author Lan, Guangchen
Xiong, Lian
Zhou, Xin
Cui, Hejie
Zhang, Yuwei
Li, Mao
Shi, Zhenyu
Fetahu, Besnik
Li, Lihong
Li, Xian
author_facet Lan, Guangchen
Xiong, Lian
Zhou, Xin
Cui, Hejie
Zhang, Yuwei
Li, Mao
Shi, Zhenyu
Fetahu, Besnik
Li, Lihong
Li, Xian
contents Reinforcement Learning with Rubric Rewards (RLRR) is a framework that extends conventional reinforcement learning from human feedback (RLHF) and verifiable rewards (RLVR) by replacing scalar preference signals with structured, multi-dimensional, contextual rubric-based evaluations. However, existing approaches in RLRR are limited to linearly compressing vector rewards into a scalar reward with a fixed weightings, which is sensitive to artificial score design and fails to capture correlations among reward dimensions. To overcome the limitations of reward aggregation, this work proposes Alternating Reinforcement Learning with Rubric Rewards (ARL-RR), a framework that eliminates the need for a fixed scalarization by optimizing one semantic rubric meta-class at a time. Theoretically, we show that reward aggregation induces a variance contraction effect, which helps explain the performance gains. We further introduce a lightweight, search-based adaptation procedure that selects the next meta-class dynamically based on task performance, enabling the policy to emphasize critical objectives and thereby improve the model performance. Empirically, our experiments on the HealthBench dataset with experts annotations demonstrate that ARL-RR uniformly outperforms scalarized methods in both model performance and training efficiency across different model scales (1.7B, 4B, 8B, and 14B).
format Preprint
id arxiv_https___arxiv_org_abs_2603_15646
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
Lan, Guangchen
Xiong, Lian
Zhou, Xin
Cui, Hejie
Zhang, Yuwei
Li, Mao
Shi, Zhenyu
Fetahu, Besnik
Li, Lihong
Li, Xian
Machine Learning
Artificial Intelligence
Computation and Language
I.2.6; I.2.7
Reinforcement Learning with Rubric Rewards (RLRR) is a framework that extends conventional reinforcement learning from human feedback (RLHF) and verifiable rewards (RLVR) by replacing scalar preference signals with structured, multi-dimensional, contextual rubric-based evaluations. However, existing approaches in RLRR are limited to linearly compressing vector rewards into a scalar reward with a fixed weightings, which is sensitive to artificial score design and fails to capture correlations among reward dimensions. To overcome the limitations of reward aggregation, this work proposes Alternating Reinforcement Learning with Rubric Rewards (ARL-RR), a framework that eliminates the need for a fixed scalarization by optimizing one semantic rubric meta-class at a time. Theoretically, we show that reward aggregation induces a variance contraction effect, which helps explain the performance gains. We further introduce a lightweight, search-based adaptation procedure that selects the next meta-class dynamically based on task performance, enabling the policy to emphasize critical objectives and thereby improve the model performance. Empirically, our experiments on the HealthBench dataset with experts annotations demonstrate that ARL-RR uniformly outperforms scalarized methods in both model performance and training efficiency across different model scales (1.7B, 4B, 8B, and 14B).
title Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
topic Machine Learning
Artificial Intelligence
Computation and Language
I.2.6; I.2.7
url https://arxiv.org/abs/2603.15646