Factored Causal Representation Learning for Robust Reward Modeling in RLHF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yupei, Yang, Lin, Deng, Wanxi, Qu, Lin, Feng, Fan, Huang, Biwei, Tu, Shikui, Xu, Lei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911689946955776
author Yang, Yupei
Yang, Lin
Deng, Wanxi
Qu, Lin
Feng, Fan
Huang, Biwei
Tu, Shikui
Xu, Lei
author_facet Yang, Yupei
Yang, Lin
Deng, Wanxi
Qu, Lin
Feng, Fan
Huang, Biwei
Tu, Shikui
Xu, Lei
contents A reliable reward model is essential for aligning large language models with human preferences through reinforcement learning from human feedback. However, standard reward models are susceptible to spurious features that are not causally related to human labels. This can lead to reward hacking, where high predicted reward does not translate into better behavior. In this work, we address this problem from a causal perspective by proposing a factored representation learning framework that decomposes the model's contextual embedding into (1) causal factors that are sufficient for reward prediction and (2) non-causal factors that capture reward-irrelevant attributes such as length or sycophantic bias. The reward head is then constrained to depend only on the causal component. In addition, we introduce an adversarial head trained to predict reward from the non-causal factors, while applying gradient reversal to discourage them from encoding reward-relevant information. Experiments on both mathematical and dialogue tasks demonstrate that our method learns more robust reward models and consistently improves downstream RLHF performance over state-of-the-art baselines. Analyses on length and sycophantic bias further validate the effectiveness of our method in mitigating reward hacking behaviors.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21350
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Yang, Yupei
Yang, Lin
Deng, Wanxi
Qu, Lin
Feng, Fan
Huang, Biwei
Tu, Shikui
Xu, Lei
Machine Learning
A reliable reward model is essential for aligning large language models with human preferences through reinforcement learning from human feedback. However, standard reward models are susceptible to spurious features that are not causally related to human labels. This can lead to reward hacking, where high predicted reward does not translate into better behavior. In this work, we address this problem from a causal perspective by proposing a factored representation learning framework that decomposes the model's contextual embedding into (1) causal factors that are sufficient for reward prediction and (2) non-causal factors that capture reward-irrelevant attributes such as length or sycophantic bias. The reward head is then constrained to depend only on the causal component. In addition, we introduce an adversarial head trained to predict reward from the non-causal factors, while applying gradient reversal to discourage them from encoding reward-relevant information. Experiments on both mathematical and dialogue tasks demonstrate that our method learns more robust reward models and consistently improves downstream RLHF performance over state-of-the-art baselines. Analyses on length and sycophantic bias further validate the effectiveness of our method in mitigating reward hacking behaviors.
title Factored Causal Representation Learning for Robust Reward Modeling in RLHF
topic Machine Learning
url https://arxiv.org/abs/2601.21350