How LLMs Are Persuaded: A Few Attention Heads, Rerouted

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Xiangkun, Kong, Lingkai, Zhang, Aoqi, Zeng, Liang, Wang, Tonghan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913107849248768
author Sun, Xiangkun
Kong, Lingkai
Zhang, Aoqi
Zeng, Liang
Wang, Tonghan
author_facet Sun, Xiangkun
Kong, Lingkai
Zhang, Aoqi
Zeng, Liang
Wang, Tonghan
contents Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A small set of mid-layer attention heads almost entirely determines the model's answer. These heads write answer options into a low-dimensional polyhedron, with options occupying distinct vertices. Persuasion does not blur belief or merely reduce confidence; it causes a discrete latent jump from the correct-answer vertex to the persuasion-target vertex. We show that decision heads are not reasoning over evidence. Instead, they copy whichever option token their attention selects. Persuasion works by redirecting attention. We isolate a rank-one evidence-routing feature that controls the route. Directly modifying this feature steers the model's choice, and removing it blocks persuasion. We then trace the feature back to a band of shallower attention heads that build it from persuasive keywords in the input. Every step is validated by intervention. This mechanism appears across open-source LLMs and realistic poisoning scenarios such as Generative Engine Optimization, revealing persuasion as a narrow, monitorable circuit.
format Preprint
id arxiv_https___arxiv_org_abs_2605_09314
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Sun, Xiangkun
Kong, Lingkai
Zhang, Aoqi
Zeng, Liang
Wang, Tonghan
Artificial Intelligence
I.2.7
Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A small set of mid-layer attention heads almost entirely determines the model's answer. These heads write answer options into a low-dimensional polyhedron, with options occupying distinct vertices. Persuasion does not blur belief or merely reduce confidence; it causes a discrete latent jump from the correct-answer vertex to the persuasion-target vertex. We show that decision heads are not reasoning over evidence. Instead, they copy whichever option token their attention selects. Persuasion works by redirecting attention. We isolate a rank-one evidence-routing feature that controls the route. Directly modifying this feature steers the model's choice, and removing it blocks persuasion. We then trace the feature back to a band of shallower attention heads that build it from persuasive keywords in the input. Every step is validated by intervention. This mechanism appears across open-source LLMs and realistic poisoning scenarios such as Generative Engine Optimization, revealing persuasion as a narrow, monitorable circuit.
title How LLMs Are Persuaded: A Few Attention Heads, Rerouted
topic Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2605.09314