Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yu, Zhao, Zihua, Huan, Zhaoxin, Gu, Wanli, Hong, Feng, Ge, Xinmu, Yuan, Lin, Wu, Weichang, Hu, Qiang, Zhang, Xiaolu, Zhou, Jun, Yao, Jiangchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916047095857152
author Huang, Yu
Zhao, Zihua
Huan, Zhaoxin
Gu, Wanli
Hong, Feng
Ge, Xinmu
Yuan, Lin
Wu, Weichang
Hu, Qiang
Zhang, Xiaolu
Zhou, Jun
Yao, Jiangchao
author_facet Huang, Yu
Zhao, Zihua
Huan, Zhaoxin
Gu, Wanli
Hong, Feng
Ge, Xinmu
Yuan, Lin
Wu, Weichang
Hu, Qiang
Zhang, Xiaolu
Zhou, Jun
Yao, Jiangchao
contents The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradigm is the imbalanced reward polarization along different rubric dimensions. Under this bottleneck, even if LLMs achieve relatively high rewards after training, they may still exhibit severe deficiencies in certain dimensions, leading to a direct deterioration in user experience. To address this problem, we propose Focal Reward, a novel objective to automatically balance the training of reinforcement learning under rubric-based rewards. Specifically, we first leverage an inverse reward projection mechanism to estimate the saturation degree of each criterion in the rubric, which forms the basis to calibrate the reward direction. Then, the final objective is designed with an automatically reweighting coefficient for each criterion to achieve the fine-grained balancing. Extensive experiments across three model scales and six benchmarks demonstrate that our Focal Reward method outperforms the strongest static aggregation baseline in all 18 model-benchmark comparisons. Rollout, mechanism, and ablation analyses further show that these gains arise from online, saturation-aware reallocation toward rubrics that still have room for improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_26579
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
Huang, Yu
Zhao, Zihua
Huan, Zhaoxin
Gu, Wanli
Hong, Feng
Ge, Xinmu
Yuan, Lin
Wu, Weichang
Hu, Qiang
Zhang, Xiaolu
Zhou, Jun
Yao, Jiangchao
Machine Learning
The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradigm is the imbalanced reward polarization along different rubric dimensions. Under this bottleneck, even if LLMs achieve relatively high rewards after training, they may still exhibit severe deficiencies in certain dimensions, leading to a direct deterioration in user experience. To address this problem, we propose Focal Reward, a novel objective to automatically balance the training of reinforcement learning under rubric-based rewards. Specifically, we first leverage an inverse reward projection mechanism to estimate the saturation degree of each criterion in the rubric, which forms the basis to calibrate the reward direction. Then, the final objective is designed with an automatically reweighting coefficient for each criterion to achieve the fine-grained balancing. Extensive experiments across three model scales and six benchmarks demonstrate that our Focal Reward method outperforms the strongest static aggregation baseline in all 18 model-benchmark comparisons. Rollout, mechanism, and ablation analyses further show that these gains arise from online, saturation-aware reallocation toward rubrics that still have room for improvement.
title Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
topic Machine Learning
url https://arxiv.org/abs/2605.26579