Mitigating Reward Hacking in RLHF via Advantage Sign Robustness

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ono, Shinnosuke, Ackermann, Johannes, Nishimori, Soichiro, Ishida, Takashi, Sugiyama, Masashi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908934878527488
author Ono, Shinnosuke
Ackermann, Johannes
Nishimori, Soichiro
Ishida, Takashi
Sugiyama, Masashi
author_facet Ono, Shinnosuke
Ackermann, Johannes
Nishimori, Soichiro
Ishida, Takashi
Sugiyama, Masashi
contents Reward models (RMs) used in reinforcement learning from human feedback (RLHF) are vulnerable to reward hacking: as the policy maximizes a learned proxy reward, true quality plateaus or degrades. We make the assumption that reward hacking is often caused by flipped advantage signs: instead of reducing the likelihood of a bad response, a flipped sign causes the update to increase it. By considering an adversarial perturbation in the RM parameter space, we can derive a certified sign-preservation radius, which is the smallest perturbation that can flip the advantage sign during policy optimization. Based on this formulation, we propose Sign-Certified Policy Optimization (SignCert-PO), down-weighting non-robust completions in the policy gradient update. Unlike prior approaches that require multiple RMs or access to the RM training data, SignCert-PO is lightweight and operates purely at the policy optimization stage using only the RM parameters and on-policy completions. On TL;DR summarization and AlpacaFarm benchmarks, SignCert-PO consistently achieves a better win rate than baselines and reduces reward hacking.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02986
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
Ono, Shinnosuke
Ackermann, Johannes
Nishimori, Soichiro
Ishida, Takashi
Sugiyama, Masashi
Machine Learning
Artificial Intelligence
Computation and Language
Reward models (RMs) used in reinforcement learning from human feedback (RLHF) are vulnerable to reward hacking: as the policy maximizes a learned proxy reward, true quality plateaus or degrades. We make the assumption that reward hacking is often caused by flipped advantage signs: instead of reducing the likelihood of a bad response, a flipped sign causes the update to increase it. By considering an adversarial perturbation in the RM parameter space, we can derive a certified sign-preservation radius, which is the smallest perturbation that can flip the advantage sign during policy optimization. Based on this formulation, we propose Sign-Certified Policy Optimization (SignCert-PO), down-weighting non-robust completions in the policy gradient update. Unlike prior approaches that require multiple RMs or access to the RM training data, SignCert-PO is lightweight and operates purely at the policy optimization stage using only the RM parameters and on-policy completions. On TL;DR summarization and AlpacaFarm benchmarks, SignCert-PO consistently achieves a better win rate than baselines and reduces reward hacking.
title Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.02986