Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914412564054016 |
|---|---|
| author | Vennemeyer, Daniel Duong, Phan Anh Zhan, Tiffany Jiang, Tianyu |
| author_facet | Vennemeyer, Daniel Duong, Phan Anh Zhan, Tiffany Jiang, Tianyu |
| contents | Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We decompose sycophancy into sycophantic agreement and sycophantic praise, contrasting both with genuine agreement. Using difference-in-means directions, activation additions, and subspace geometry across multiple models and datasets, we show that: (1) the three behaviors are encoded along distinct linear directions in latent space; (2) each behavior can be independently amplified or suppressed without affecting the others; and (3) their representational structure is consistent across model families and scales. These results suggest that sycophantic behaviors correspond to distinct, independently steerable representations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_21305 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs Vennemeyer, Daniel Duong, Phan Anh Zhan, Tiffany Jiang, Tianyu Computation and Language Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We decompose sycophancy into sycophantic agreement and sycophantic praise, contrasting both with genuine agreement. Using difference-in-means directions, activation additions, and subspace geometry across multiple models and datasets, we show that: (1) the three behaviors are encoded along distinct linear directions in latent space; (2) each behavior can be independently amplified or suppressed without affecting the others; and (3) their representational structure is consistent across model families and scales. These results suggest that sycophantic behaviors correspond to distinct, independently steerable representations. |
| title | Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2509.21305 |