ELEPHANT: Measuring and understanding social sycophancy in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cheng, Myra, Yu, Sunny, Lee, Cinoo, Khadpe, Pranav, Ibrahim, Lujain, Jurafsky, Dan
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917381109972992
author Cheng, Myra
Yu, Sunny
Lee, Cinoo
Khadpe, Pranav
Ibrahim, Lujain
Jurafsky, Dan
author_facet Cheng, Myra
Yu, Sunny
Lee, Cinoo
Khadpe, Pranav
Ibrahim, Lujain
Jurafsky, Dan
contents LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth. This fails to capture broader forms of sycophancy such as affirming a user's self-image or other implicit beliefs. To address this gap, we introduce social sycophancy, characterizing sycophancy as excessive preservation of a user's face (their desired self-image), and present ELEPHANT, a benchmark for measuring social sycophancy in an LLM. Applying our benchmark to 11 models, we show that LLMs consistently exhibit high rates of social sycophancy: on average, they preserve user's face 45 percentage points more than humans in general advice queries and in queries describing clear user wrongdoing (from Reddit's r/AmITheAsshole). Furthermore, when prompted with perspectives from either side of a moral conflict, LLMs affirm both sides (depending on whichever side the user adopts) in 48% of cases--telling both the at-fault party and the wronged party that they are not wrong--rather than adhering to a consistent moral or value judgment. We further show that social sycophancy is rewarded in preference datasets, and that while existing mitigation strategies for sycophancy are limited in effectiveness, model-based steering shows promise for mitigating these behaviors. Our work provides theoretical grounding and an empirical benchmark for understanding and addressing sycophancy in the open-ended contexts that characterize the vast majority of LLM use cases.
format Preprint
id arxiv_https___arxiv_org_abs_2505_13995
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ELEPHANT: Measuring and understanding social sycophancy in LLMs
Cheng, Myra
Yu, Sunny
Lee, Cinoo
Khadpe, Pranav
Ibrahim, Lujain
Jurafsky, Dan
Computation and Language
Artificial Intelligence
Computers and Society
LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth. This fails to capture broader forms of sycophancy such as affirming a user's self-image or other implicit beliefs. To address this gap, we introduce social sycophancy, characterizing sycophancy as excessive preservation of a user's face (their desired self-image), and present ELEPHANT, a benchmark for measuring social sycophancy in an LLM. Applying our benchmark to 11 models, we show that LLMs consistently exhibit high rates of social sycophancy: on average, they preserve user's face 45 percentage points more than humans in general advice queries and in queries describing clear user wrongdoing (from Reddit's r/AmITheAsshole). Furthermore, when prompted with perspectives from either side of a moral conflict, LLMs affirm both sides (depending on whichever side the user adopts) in 48% of cases--telling both the at-fault party and the wronged party that they are not wrong--rather than adhering to a consistent moral or value judgment. We further show that social sycophancy is rewarded in preference datasets, and that while existing mitigation strategies for sycophancy are limited in effectiveness, model-based steering shows promise for mitigating these behaviors. Our work provides theoretical grounding and an empirical benchmark for understanding and addressing sycophancy in the open-ended contexts that characterize the vast majority of LLM use cases.
title ELEPHANT: Measuring and understanding social sycophancy in LLMs
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2505.13995