Reward Model Interpretability via Optimal and Pessimal Tokens

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Christian, Brian, Kirk, Hannah Rose, Thompson, Jessica A. F., Summerfield, Christopher, Dumbalska, Tsvetomira
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917242714718208
author Christian, Brian
Kirk, Hannah Rose
Thompson, Jessica A. F.
Summerfield, Christopher
Dumbalska, Tsvetomira
author_facet Christian, Brian
Kirk, Hannah Rose
Thompson, Jessica A. F.
Summerfield, Christopher
Dumbalska, Tsvetomira
contents Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine-tuning generative models. However, the reward models themselves -- which directly encode human value judgments by turning prompt-response pairs into scalar rewards -- remain relatively understudied. We present a novel approach to reward model interpretability through exhaustive analysis of their responses across their entire vocabulary space. By examining how different reward models score every possible single-token response to value-laden prompts, we uncover several striking findings: (i) substantial heterogeneity between models trained on similar objectives, (ii) systematic asymmetries in how models encode high- vs low-scoring tokens, (iii) significant sensitivity to prompt framing that mirrors human cognitive biases, and (iv) overvaluation of more frequent tokens. We demonstrate these effects across ten recent open-source reward models of varying parameter counts and architectures. Our results challenge assumptions about the interchangeability of reward models, as well as their suitability as proxies of complex and context-dependent human values. We find that these models can encode concerning biases toward certain identity groups, which may emerge as unintended consequences of harmlessness training -- distortions that risk propagating through the downstream large language models now deployed to millions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07326
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reward Model Interpretability via Optimal and Pessimal Tokens
Christian, Brian
Kirk, Hannah Rose
Thompson, Jessica A. F.
Summerfield, Christopher
Dumbalska, Tsvetomira
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
I.2.6; I.2.7; H.5.2; J.4; K.4.2
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine-tuning generative models. However, the reward models themselves -- which directly encode human value judgments by turning prompt-response pairs into scalar rewards -- remain relatively understudied. We present a novel approach to reward model interpretability through exhaustive analysis of their responses across their entire vocabulary space. By examining how different reward models score every possible single-token response to value-laden prompts, we uncover several striking findings: (i) substantial heterogeneity between models trained on similar objectives, (ii) systematic asymmetries in how models encode high- vs low-scoring tokens, (iii) significant sensitivity to prompt framing that mirrors human cognitive biases, and (iv) overvaluation of more frequent tokens. We demonstrate these effects across ten recent open-source reward models of varying parameter counts and architectures. Our results challenge assumptions about the interchangeability of reward models, as well as their suitability as proxies of complex and context-dependent human values. We find that these models can encode concerning biases toward certain identity groups, which may emerge as unintended consequences of harmlessness training -- distortions that risk propagating through the downstream large language models now deployed to millions.
title Reward Model Interpretability via Optimal and Pessimal Tokens
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
I.2.6; I.2.7; H.5.2; J.4; K.4.2
url https://arxiv.org/abs/2506.07326