When Are Two RLHF Objectives the Same?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Gaikwad, Madhava
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911423466045440
author Gaikwad, Madhava
author_facet Gaikwad, Madhava
contents The preference optimization literature contains many proposed objectives, often presented as distinct improvements. We introduce Opal, a canonicalization algorithm that determines whether two preference objectives are algebraically equivalent by producing either a canonical form or a concrete witness of non-equivalence. Applying Opal reveals that many widely used methods optimize the same underlying objective, while others are provably distinct. For example, batch normalization can cause the same response pair to receive different gradients depending on batch composition. We identify a small set of structural mechanisms that give rise to genuinely different objectives; most remaining differences are reparameterizations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Are Two RLHF Objectives the Same?
Gaikwad, Madhava
Machine Learning
Artificial Intelligence
Computation and Language
68T05, 68T07, 68Q32, 62H30, 62F15, 90C30
I.2.6; I.2.7; I.2.8; G.3; G.1.6
The preference optimization literature contains many proposed objectives, often presented as distinct improvements. We introduce Opal, a canonicalization algorithm that determines whether two preference objectives are algebraically equivalent by producing either a canonical form or a concrete witness of non-equivalence. Applying Opal reveals that many widely used methods optimize the same underlying objective, while others are provably distinct. For example, batch normalization can cause the same response pair to receive different gradients depending on batch composition. We identify a small set of structural mechanisms that give rise to genuinely different objectives; most remaining differences are reparameterizations.
title When Are Two RLHF Objectives the Same?
topic Machine Learning
Artificial Intelligence
Computation and Language
68T05, 68T07, 68Q32, 62H30, 62F15, 90C30
I.2.6; I.2.7; I.2.8; G.3; G.1.6
url https://arxiv.org/abs/2509.11298