Sycophancy Hides Linearly in the Attention Heads

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Genadi, Rifo, Nwadike, Munachiso, Mukhituly, Nurdaulet, Alquabeh, Hilal, Hiraoka, Tatsuya, Inui, Kentaro
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914275440721920
author Genadi, Rifo
Nwadike, Munachiso
Mukhituly, Nurdaulet
Alquabeh, Hilal
Hiraoka, Tatsuya
Inui, Kentaro
author_facet Genadi, Rifo
Nwadike, Munachiso
Mukhituly, Nurdaulet
Alquabeh, Hilal
Hiraoka, Tatsuya
Inui, Kentaro
contents We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer perceptron (MLP), and attention layers to analyze where these signals emerge. Although separability appears in the residual stream and MLPs, steering using these probes is most effective in a sparse subset of middle-layer attention heads. Using TruthfulQA as the base dataset, we find that probes trained on it transfer effectively to other factual QA benchmarks. Furthermore, comparing our discovered direction to previously identified "truthful" directions reveals limited overlap, suggesting that factual accuracy, and deference resistance, arise from related but distinct mechanisms. Attention-pattern analysis further indicates that the influential heads attend disproportionately to expressions of user doubt, contributing to sycophantic shifts. Overall, these findings suggest that sycophancy can be mitigated through simple, targeted linear interventions that exploit the internal geometry of attention activations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16644
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Sycophancy Hides Linearly in the Attention Heads
Genadi, Rifo
Nwadike, Munachiso
Mukhituly, Nurdaulet
Alquabeh, Hilal
Hiraoka, Tatsuya
Inui, Kentaro
Computation and Language
Artificial Intelligence
We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer perceptron (MLP), and attention layers to analyze where these signals emerge. Although separability appears in the residual stream and MLPs, steering using these probes is most effective in a sparse subset of middle-layer attention heads. Using TruthfulQA as the base dataset, we find that probes trained on it transfer effectively to other factual QA benchmarks. Furthermore, comparing our discovered direction to previously identified "truthful" directions reveals limited overlap, suggesting that factual accuracy, and deference resistance, arise from related but distinct mechanisms. Attention-pattern analysis further indicates that the influential heads attend disproportionately to expressions of user doubt, contributing to sycophantic shifts. Overall, these findings suggest that sycophancy can be mitigated through simple, targeted linear interventions that exploit the internal geometry of attention activations.
title Sycophancy Hides Linearly in the Attention Heads
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.16644