Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yuefei, Quan, Yihao, Lin, Xiaodong, Tang, Ruixiang
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915947016617984
author Chen, Yuefei
Quan, Yihao
Lin, Xiaodong
Tang, Ruixiang
author_facet Chen, Yuefei
Quan, Yihao
Lin, Xiaodong
Tang, Ruixiang
contents LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names fail far more often than other fields across all models and settings. Citation style has no measurable effect, while reasoning-oriented distillation degrades recall. Probes trained on one field transfer at near-chance levels to the others, suggesting that hallucination signals do not generalize across fields. Building on this finding, we apply elastic-net regularization with stability selection to neuron-level CETT values of Qwen2.5-32B-Instruct and identify a sparse set of field-specific hallucination neurons (FH-neurons). Causal intervention further confirms their role: amplifying these neurons increases hallucination, while suppressing them improves performance across fields, with larger gains in some fields. These results suggest a lightweight approach to detecting and mitigating citation hallucination using internal model signals alone.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18880
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
Chen, Yuefei
Quan, Yihao
Lin, Xiaodong
Tang, Ruixiang
Computation and Language
Artificial Intelligence
LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names fail far more often than other fields across all models and settings. Citation style has no measurable effect, while reasoning-oriented distillation degrades recall. Probes trained on one field transfer at near-chance levels to the others, suggesting that hallucination signals do not generalize across fields. Building on this finding, we apply elastic-net regularization with stability selection to neuron-level CETT values of Qwen2.5-32B-Instruct and identify a sparse set of field-specific hallucination neurons (FH-neurons). Causal intervention further confirms their role: amplifying these neurons increases hallucination, while suppressing them improves performance across fields, with larger gains in some fields. These results suggest a lightweight approach to detecting and mitigating citation hallucination using internal model signals alone.
title Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.18880