When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Tong, Bai, Andrew, Ban, Yuanhao, Hong, Yunqi, Li, Haoyu, Hsieh, Cho-jui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917237989834752
author Xie, Tong
Bai, Andrew
Ban, Yuanhao
Hong, Yunqi
Li, Haoyu
Hsieh, Cho-jui
author_facet Xie, Tong
Bai, Andrew
Ban, Yuanhao
Hong, Yunqi
Li, Haoyu
Hsieh, Cho-jui
contents Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which learns from pairwise data consisting of chosen and rejected responses. In this work, we analyze the per-sample gradient of BT-loss and show spurious learning signals due to representation distance. In particular, BT gradient norm scales with two distinct components: (1) prediction error, reflected by the difference in predicted rewards between chosen and rejected responses, and critically, (2) representation distance between the pair measured in the output space of the final layer. While the first term captures the intended training signal, the second term can significantly impact the update magnitude and misalign learning. Specifically, pairs with small representation distance often receive vanishingly weak updates, even when misranked, while pairs with large distance receive disproportionately strong updates. This leads to gradients from large-distance pairs to overshadow those from small-distance pairs, where fine-grained distinctions are especially important. To overcome this limitation, we propose NormBT, an adaptive pair-wise normalization scheme that rescales updates to balance representation-driven effects and focuses learning signals on prediction error. NormBT is a lightweight, drop-in modification to BT loss with negligible overhead. Across various LLM backbones and datasets, NormBT improves reward model performance consistently, with notable gains of over 5% on the Reasoning category of RewardBench, which contains numerous fine-grained pairs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_06343
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Xie, Tong
Bai, Andrew
Ban, Yuanhao
Hong, Yunqi
Li, Haoyu
Hsieh, Cho-jui
Machine Learning
Artificial Intelligence
Computation and Language
Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which learns from pairwise data consisting of chosen and rejected responses. In this work, we analyze the per-sample gradient of BT-loss and show spurious learning signals due to representation distance. In particular, BT gradient norm scales with two distinct components: (1) prediction error, reflected by the difference in predicted rewards between chosen and rejected responses, and critically, (2) representation distance between the pair measured in the output space of the final layer. While the first term captures the intended training signal, the second term can significantly impact the update magnitude and misalign learning. Specifically, pairs with small representation distance often receive vanishingly weak updates, even when misranked, while pairs with large distance receive disproportionately strong updates. This leads to gradients from large-distance pairs to overshadow those from small-distance pairs, where fine-grained distinctions are especially important. To overcome this limitation, we propose NormBT, an adaptive pair-wise normalization scheme that rescales updates to balance representation-driven effects and focuses learning signals on prediction error. NormBT is a lightweight, drop-in modification to BT loss with negligible overhead. Across various LLM backbones and datasets, NormBT improves reward model performance consistently, with notable gains of over 5% on the Reasoning category of RewardBench, which contains numerous fine-grained pairs.
title When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2512.06343