Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ghasemi, Majid, Crowley, Mark
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914314611326976
author Ghasemi, Majid
Crowley, Mark
author_facet Ghasemi, Majid
Crowley, Mark
contents Contemporary AI alignment strategies rely on a fragile premise: that human feedback, while noisy, remains a fundamentally truthful signal. In this paper, we identify this assumption as Dogma 4 of Reinforcement Learning (RL). We demonstrate that while this dogma holds in static environments, it fails in social settings where evaluators may be sycophantic, lazy, or adversarial. We prove that under Dogma 4, standard RL agents suffer from what we call Objective Decoupling, a structural failure mode where the agent's learned objective permanently separates from the latent ground truth, guaranteeing convergence to misalignment. To resolve this, we propose Epistemic Source Alignment (ESA). Unlike standard robust methods that rely on statistical consensus (trusting the majority), ESA utilizes sparse safety axioms to judge the source of the feedback rather than the signal itself. We prove that this "judging the judges" mechanism guarantees convergence to the true objective, even when a majority of evaluators are biased. Empirically, we show that while traditional consensus methods fail under majority collusion, our approach successfully recovers the optimal policy.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08092
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities
Ghasemi, Majid
Crowley, Mark
Artificial Intelligence
Emerging Technologies
Contemporary AI alignment strategies rely on a fragile premise: that human feedback, while noisy, remains a fundamentally truthful signal. In this paper, we identify this assumption as Dogma 4 of Reinforcement Learning (RL). We demonstrate that while this dogma holds in static environments, it fails in social settings where evaluators may be sycophantic, lazy, or adversarial. We prove that under Dogma 4, standard RL agents suffer from what we call Objective Decoupling, a structural failure mode where the agent's learned objective permanently separates from the latent ground truth, guaranteeing convergence to misalignment. To resolve this, we propose Epistemic Source Alignment (ESA). Unlike standard robust methods that rely on statistical consensus (trusting the majority), ESA utilizes sparse safety axioms to judge the source of the feedback rather than the signal itself. We prove that this "judging the judges" mechanism guarantees convergence to the true objective, even when a majority of evaluators are biased. Empirically, we show that while traditional consensus methods fail under majority collusion, our approach successfully recovers the optimal policy.
title Objective Decoupling in Social Reinforcement Learning: Recovering Ground Truth from Sycophantic Majorities
topic Artificial Intelligence
Emerging Technologies
url https://arxiv.org/abs/2602.08092