Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vennemeyer, Daniel, Duong, Phan Anh, Zhan, Tiffany, Jiang, Tianyu
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914412564054016
author Vennemeyer, Daniel
Duong, Phan Anh
Zhan, Tiffany
Jiang, Tianyu
author_facet Vennemeyer, Daniel
Duong, Phan Anh
Zhan, Tiffany
Jiang, Tianyu
contents Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We decompose sycophancy into sycophantic agreement and sycophantic praise, contrasting both with genuine agreement. Using difference-in-means directions, activation additions, and subspace geometry across multiple models and datasets, we show that: (1) the three behaviors are encoded along distinct linear directions in latent space; (2) each behavior can be independently amplified or suppressed without affecting the others; and (3) their representational structure is consistent across model families and scales. These results suggest that sycophantic behaviors correspond to distinct, independently steerable representations.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
Vennemeyer, Daniel
Duong, Phan Anh
Zhan, Tiffany
Jiang, Tianyu
Computation and Language
Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We decompose sycophancy into sycophantic agreement and sycophantic praise, contrasting both with genuine agreement. Using difference-in-means directions, activation additions, and subspace geometry across multiple models and datasets, we show that: (1) the three behaviors are encoded along distinct linear directions in latent space; (2) each behavior can be independently amplified or suppressed without affecting the others; and (3) their representational structure is consistent across model families and scales. These results suggest that sycophantic behaviors correspond to distinct, independently steerable representations.
title Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
topic Computation and Language
url https://arxiv.org/abs/2509.21305