RLHF May Not Reflect Genuine Preferences

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghafouri, Bijean, Choi, Eun Cheol, Dey, Priyanka, Ferrara, Emilio
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914618735067136
author Ghafouri, Bijean
Choi, Eun Cheol
Dey, Priyanka
Ferrara, Emilio
author_facet Ghafouri, Bijean
Choi, Eun Cheol
Dey, Priyanka
Ferrara, Emilio
contents Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Importantly, these failures are common for the judgments on values that matter most for AI alignment. We argue that measurement validity is logically prior to preference aggregation. Before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. We organize annotation responses along a spectrum, from non-attitudes (no signal) to genuine preferences (full signal), and develop diagnostics that locate responses on this spectrum. In two RLHF datasets, we show that inconsistency is systematic and directionally biased. Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale. As such, much of the current RLHF practice models noise as signal and elicitation artifacts as human values.
format Preprint
id arxiv_https___arxiv_org_abs_2604_03238
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RLHF May Not Reflect Genuine Preferences
Ghafouri, Bijean
Choi, Eun Cheol
Dey, Priyanka
Ferrara, Emilio
Human-Computer Interaction
Reinforcement Learning from Human Feedback (RLHF) assumes that annotation responses reflect genuine human preferences. They often do not. Behavioral scientists have documented for sixty years that people produce responses without holding genuine opinions, construct preferences on the spot from contextual cues, and interpret identical questions differently. Importantly, these failures are common for the judgments on values that matter most for AI alignment. We argue that measurement validity is logically prior to preference aggregation. Before asking how to combine annotations, the field must ask whether the responses being combined are preferences at all. We organize annotation responses along a spectrum, from non-attitudes (no signal) to genuine preferences (full signal), and develop diagnostics that locate responses on this spectrum. In two RLHF datasets, we show that inconsistency is systematic and directionally biased. Filtering high-inconsistency annotators flips majority harm classifications for 18.6% of prompts and shifts mean ratings by over 13 points on a 100-point scale. As such, much of the current RLHF practice models noise as signal and elicitation artifacts as human values.
title RLHF May Not Reflect Genuine Preferences
topic Human-Computer Interaction
url https://arxiv.org/abs/2604.03238