Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pandey, Sanskar, Chopra, Ruhaan, Puniya, Angkul, Pal, Sohom
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917503953797120
author Pandey, Sanskar
Chopra, Ruhaan
Puniya, Angkul
Pal, Sohom
author_facet Pandey, Sanskar
Chopra, Ruhaan
Puniya, Angkul
Pal, Sohom
contents Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as a preference for user agreement over principled reasoning. We introduce Beacon, a single-turn forced-choice benchmark that isolates this bias independent of conversational context, enabling precise measurement of the tension between factual accuracy and submissive bias. Evaluations across twelve state-of-the-art models reveal that sycophancy decomposes into stable linguistic and affective sub-biases, each scaling with model capacity. We further propose prompt-level and activation-level interventions that modulate these biases in opposing directions, exposing the internal geometry of alignment as a dynamic manifold between truthfulness and socially compliant judgment. Beacon reframes sycophancy as a measurable form of normative misgeneralization, providing a reproducible foundation for studying and mitigating alignment drift in large-scale generative systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16727
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
Pandey, Sanskar
Chopra, Ruhaan
Puniya, Angkul
Pal, Sohom
Computation and Language
Artificial Intelligence
Large language models internalize a structural trade-off between truthfulness and obsequious flattery, emerging from reward optimization that conflates helpfulness with polite submission. This latent bias, known as sycophancy, manifests as a preference for user agreement over principled reasoning. We introduce Beacon, a single-turn forced-choice benchmark that isolates this bias independent of conversational context, enabling precise measurement of the tension between factual accuracy and submissive bias. Evaluations across twelve state-of-the-art models reveal that sycophancy decomposes into stable linguistic and affective sub-biases, each scaling with model capacity. We further propose prompt-level and activation-level interventions that modulate these biases in opposing directions, exposing the internal geometry of alignment as a dynamic manifold between truthfulness and socially compliant judgment. Beacon reframes sycophancy as a measurable form of normative misgeneralization, providing a reproducible foundation for studying and mitigating alignment drift in large-scale generative systems.
title Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.16727