Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912936031682560 |
|---|---|
| author | Zhang, Junkai Wang, Zihao Gui, Lin Sathyendra, Swarnashree Mysore Jeong, Jaehwan Veitch, Victor Wang, Wei He, Yunzhong Liu, Bing Jin, Lifeng |
| author_facet | Zhang, Junkai Wang, Zihao Gui, Lin Sathyendra, Swarnashree Mysore Jeong, Jaehwan Veitch, Victor Wang, Wei He, Yunzhong Liu, Bing Jin, Lifeng |
| contents | Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to reliably distinguish Excellent responses from merely Great ones. This motivate us to focus on the high-reward region. However, such tail examples are scarce under the base LLM. While off-policy exemplars (e.g. from stronger models or rewrites) are easier to obtain, naively training on them yields a misspecified reward for the policy we aim to align. To address this, we study rubric-based rewards. By design, rubrics can leverage off-policy examples while remaining insensitive to their artifacts. To elicit rubrics that capture the high-reward tail, we highlight the importance of distinguishing among great and diverse responses, and introduce a workflow to implement this idea. We empirically demonstrate that rubric-based rewards substantially mitigate reward over-optimization and deliver effective LLM post-training improvements. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21500 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training Zhang, Junkai Wang, Zihao Gui, Lin Sathyendra, Swarnashree Mysore Jeong, Jaehwan Veitch, Victor Wang, Wei He, Yunzhong Liu, Bing Jin, Lifeng Machine Learning Artificial Intelligence 68T50 I.2 Reinforcement fine-tuning (RFT) often suffers from reward over-optimization, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the inability to reliably distinguish Excellent responses from merely Great ones. This motivate us to focus on the high-reward region. However, such tail examples are scarce under the base LLM. While off-policy exemplars (e.g. from stronger models or rewrites) are easier to obtain, naively training on them yields a misspecified reward for the policy we aim to align. To address this, we study rubric-based rewards. By design, rubrics can leverage off-policy examples while remaining insensitive to their artifacts. To elicit rubrics that capture the high-reward tail, we highlight the importance of distinguishing among great and diverse responses, and introduce a workflow to implement this idea. We empirically demonstrate that rubric-based rewards substantially mitigate reward over-optimization and deliver effective LLM post-training improvements. |
| title | Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training |
| topic | Machine Learning Artificial Intelligence 68T50 I.2 |
| url | https://arxiv.org/abs/2509.21500 |