Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lamparth, Max, Fein, Daniel, Haupt, Andreas, Hussing, Marcel, Kochenderfer, Mykel J.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911730643238912
author Lamparth, Max
Fein, Daniel
Haupt, Andreas
Hussing, Marcel
Kochenderfer, Mykel J.
author_facet Lamparth, Max
Fein, Daniel
Haupt, Andreas
Hussing, Marcel
Kochenderfer, Mykel J.
contents Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. The failure is enabled by a measurement-versus-optimization gap between audit and policy-induced distributions during mitigation evaluation and policy training. We formalize mitigation outcomes into a regime taxonomy and prove that successful mitigation, bias substitution, and overcorrection produce identical observables under any audit-distribution scoring, including ranking accuracy and win-rate, even when granted oracle access to the true reward. Across published preference-learning mitigation work, no method we survey reports the evidence needed to certify successful mitigation. Augmenting evaluation with policy-induced distributions while tracking multiple biases provably closes the gap, and we translate this into actionable prescriptions for mitigation methods and benchmarks. We demonstrate bias substitution in language model RLHF, where a length penalty during GRPO training compresses responses as intended yet redirects optimization pressure onto confidence calibration, driving the policy into overconfidence while factual free-form accuracy falls. We also show a published length-debiasing operator that zeroes reward-length correlation on the audit distribution but reintroduces bias under best-of-N selection on three of four SOTA reward models, and a length-sycophancy coupling whose direction reverses under human-LLM judge disagreement.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27996
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
Lamparth, Max
Fein, Daniel
Haupt, Andreas
Hussing, Marcel
Kochenderfer, Mykel J.
Artificial Intelligence
Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto correlated proxies rather than eliminate it, a failure mode we call reward bias substitution. The failure is enabled by a measurement-versus-optimization gap between audit and policy-induced distributions during mitigation evaluation and policy training. We formalize mitigation outcomes into a regime taxonomy and prove that successful mitigation, bias substitution, and overcorrection produce identical observables under any audit-distribution scoring, including ranking accuracy and win-rate, even when granted oracle access to the true reward. Across published preference-learning mitigation work, no method we survey reports the evidence needed to certify successful mitigation. Augmenting evaluation with policy-induced distributions while tracking multiple biases provably closes the gap, and we translate this into actionable prescriptions for mitigation methods and benchmarks. We demonstrate bias substitution in language model RLHF, where a length penalty during GRPO training compresses responses as intended yet redirects optimization pressure onto confidence calibration, driving the policy into overconfidence while factual free-form accuracy falls. We also show a published length-debiasing operator that zeroes reward-length correlation on the audit distribution but reintroduces bias under best-of-N selection on three of four SOTA reward models, and a length-sycophancy coupling whose direction reverses under human-LLM judge disagreement.
title Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
topic Artificial Intelligence
url https://arxiv.org/abs/2605.27996