One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fein, Daniel, Lamparth, Max, Xiang, Violet, Kochenderfer, Mykel J., Haber, Nick
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910278331924480
author Fein, Daniel
Lamparth, Max
Xiang, Violet
Kochenderfer, Mykel J.
Haber, Nick
author_facet Fein, Daniel
Lamparth, Max
Xiang, Violet
Kochenderfer, Mykel J.
Haber, Nick
contents Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific ``styles'' and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03291
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
Fein, Daniel
Lamparth, Max
Xiang, Violet
Kochenderfer, Mykel J.
Haber, Nick
Computation and Language
Artificial Intelligence
Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, including the state-of-the-art, we find that issues persist despite prior work with respect to length, sycophancy, and overconfidence. We also discover new issues related to bias toward model-specific ``styles'' and answer-order. We categorize RM failures as tractable or resistant to linear intervention and propose a simple post-hoc intervention to mitigate low-complexity biases that arise from spurious correlations. Our proposed mechanistic reward shaping reduces targeted biases without degrading reward quality and while using minimal labeled data. The method is extensible to new biases, model-internal, and generalizes out-of-distribution.
title One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2603.03291