H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Cheng, Chen, Huimin, Xiao, Chaojun, Chen, Zhiyi, Liu, Zhiyuan, Sun, Maosong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918227324436480
author Gao, Cheng
Chen, Huimin
Xiao, Chaojun
Chen, Zhiyi
Liu, Zhiyuan
Sun, Maosong
author_facet Gao, Cheng
Chen, Huimin
Xiao, Chaojun
Chen, Zhiyi
Liu, Zhiyuan
Sun, Maosong
contents Large language models (LLMs) frequently generate hallucinations -- plausible but factually incorrect outputs -- undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored. In this paper, we conduct a systematic investigation into hallucination-associated neurons (H-Neurons) in LLMs from three perspectives: identification, behavioral impact, and origins. Regarding their identification, we demonstrate that a remarkably sparse subset of neurons (less than $0.1\%$ of total neurons) can reliably predict hallucination occurrences, with strong generalization across diverse scenarios. In terms of behavioral impact, controlled interventions reveal that these neurons are causally linked to over-compliance behaviors. Concerning their origins, we trace these neurons back to the pre-trained base models and find that these neurons remain predictive for hallucination detection, indicating they emerge during pre-training. Our findings bridge macroscopic behavioral patterns with microscopic neural mechanisms, offering insights for developing more reliable LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01797
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
Gao, Cheng
Chen, Huimin
Xiao, Chaojun
Chen, Zhiyi
Liu, Zhiyuan
Sun, Maosong
Artificial Intelligence
Computation and Language
Computers and Society
Large language models (LLMs) frequently generate hallucinations -- plausible but factually incorrect outputs -- undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored. In this paper, we conduct a systematic investigation into hallucination-associated neurons (H-Neurons) in LLMs from three perspectives: identification, behavioral impact, and origins. Regarding their identification, we demonstrate that a remarkably sparse subset of neurons (less than $0.1\%$ of total neurons) can reliably predict hallucination occurrences, with strong generalization across diverse scenarios. In terms of behavioral impact, controlled interventions reveal that these neurons are causally linked to over-compliance behaviors. Concerning their origins, we trace these neurons back to the pre-trained base models and find that these neurons remain predictive for hallucination detection, indicating they emerge during pre-training. Our findings bridge macroscopic behavioral patterns with microscopic neural mechanisms, offering insights for developing more reliable LLMs.
title H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
topic Artificial Intelligence
Computation and Language
Computers and Society
url https://arxiv.org/abs/2512.01797