Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Junkai, Wang, Zihao, Gui, Lin, Sathyendra, Swarnashree Mysore, Jeong, Jaehwan, Veitch, Victor, Wang, Wei, He, Yunzhong, Liu, Bing, Jin, Lifeng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912936031682560
author Zhang, Junkai
Wang, Zihao
Gui, Lin
Sathyendra, Swarnashree Mysore
Jeong, Jaehwan
Veitch, Victor
Wang, Wei
He, Yunzhong
Liu, Bing
Jin, Lifeng
author_facet Zhang, Junkai
Wang, Zihao
Gui, Lin
Sathyendra, Swarnashree Mysore
Jeong, Jaehwan
Veitch, Victor
Wang, Wei
He, Yunzhong
Liu, Bing
Jin, Lifeng
contents Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to reliably distinguish Excellent responses from merely Great ones. This motivate us to focus on the high-reward region. However, such tail examples are scarce under the base LLM. While off-policy exemplars (e.g. from stronger models or rewrites) are easier to obtain, naively training on them yields a misspecified reward for the policy we aim to align. To address this, we study rubric-based rewards. By design, rubrics can leverage off-policy examples while remaining insensitive to their artifacts. To elicit rubrics that capture the high-reward tail, we highlight the importance of distinguishing among great and diverse responses, and introduce a workflow to implement this idea. We empirically demonstrate that rubric-based rewards substantially mitigate reward over-optimization and deliver effective LLM post-training improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
Zhang, Junkai
Wang, Zihao
Gui, Lin
Sathyendra, Swarnashree Mysore
Jeong, Jaehwan
Veitch, Victor
Wang, Wei
He, Yunzhong
Liu, Bing
Jin, Lifeng
Machine Learning
Artificial Intelligence
68T50
I.2
Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to reliably distinguish Excellent responses from merely Great ones. This motivate us to focus on the high-reward region. However, such tail examples are scarce under the base LLM. While off-policy exemplars (e.g. from stronger models or rewrites) are easier to obtain, naively training on them yields a misspecified reward for the policy we aim to align. To address this, we study rubric-based rewards. By design, rubrics can leverage off-policy examples while remaining insensitive to their artifacts. To elicit rubrics that capture the high-reward tail, we highlight the importance of distinguishing among great and diverse responses, and introduce a workflow to implement this idea. We empirically demonstrate that rubric-based rewards substantially mitigate reward over-optimization and deliver effective LLM post-training improvements.
title Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
topic Machine Learning
Artificial Intelligence
68T50
I.2
url https://arxiv.org/abs/2509.21500