Ask don't tell: Reducing sycophancy in large language models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dubois, Magda, Ududec, Cozmin, Summerfield, Christopher, Luettgau, Lennart
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915963756085248
author Dubois, Magda
Ududec, Cozmin
Summerfield, Christopher
Luettgau, Lennart
author_facet Dubois, Magda
Ududec, Cozmin
Summerfield, Christopher
Luettgau, Lennart
contents Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While prior work has documented conversational features correlated with sycophancy, we lack a systematic understanding of what provokes or prevents AI sycophancy. Here, we present a set of controlled experimental studies where we first isolate how input framing influences sycophancy, and second, leverage these findings to develop mitigation strategies. In a nested factorial design, we compare questions to various non-questions where we vary three orthogonal factors: epistemic certainty (statement, belief, conviction), perspective (I- vs user-perspective), and affirmation vs negation. We show that (1) sycophancy is substantially higher in response to non-questions compared to questions. Additionally, we find that (2) sycophancy increases monotonically with epistemic certainty conveyed by the user, and (3) is amplified by I-perspective framing. Building on this, we show that asking a model to convert non-questions into questions before answering significantly reduces sycophancy. Importantly, this effect is stronger than a simple baseline prompt asking models "not to be sycophantic". Our work offers a practical and effective input-level mitigation that both developers and users can easily adopt.
format Preprint
id arxiv_https___arxiv_org_abs_2602_23971
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Ask don't tell: Reducing sycophancy in large language models
Dubois, Magda
Ududec, Cozmin
Summerfield, Christopher
Luettgau, Lennart
Human-Computer Interaction
Artificial Intelligence
Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While prior work has documented conversational features correlated with sycophancy, we lack a systematic understanding of what provokes or prevents AI sycophancy. Here, we present a set of controlled experimental studies where we first isolate how input framing influences sycophancy, and second, leverage these findings to develop mitigation strategies. In a nested factorial design, we compare questions to various non-questions where we vary three orthogonal factors: epistemic certainty (statement, belief, conviction), perspective (I- vs user-perspective), and affirmation vs negation. We show that (1) sycophancy is substantially higher in response to non-questions compared to questions. Additionally, we find that (2) sycophancy increases monotonically with epistemic certainty conveyed by the user, and (3) is amplified by I-perspective framing. Building on this, we show that asking a model to convert non-questions into questions before answering significantly reduces sycophancy. Importantly, this effect is stronger than a simple baseline prompt asking models "not to be sycophantic". Our work offers a practical and effective input-level mitigation that both developers and users can easily adopt.
title Ask don't tell: Reducing sycophancy in large language models
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2602.23971