When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Shaowen, Dong, Yiqi, Chang, Ruinian, Zhu, Tansheng, Sun, Yuebo, Lyu, Kaifeng, Li, Jian
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908668844310528
author Wang, Shaowen
Dong, Yiqi
Chang, Ruinian
Zhu, Tansheng
Sun, Yuebo
Lyu, Kaifeng
Li, Jian
author_facet Wang, Shaowen
Dong, Yiqi
Chang, Ruinian
Zhu, Tansheng
Sun, Yuebo
Lyu, Kaifeng
Li, Jian
contents Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and proprietary LLMs (including GPT-5), we show that existing hallucination detection methods, such as confidence-based filtering and inner-state probing, fundamentally fail in the presence of spurious correlations. Our theoretical analysis further elucidates why these statistical biases intrinsically undermine confidence-based detection techniques. Our findings thus emphasize the urgent need for new approaches explicitly designed to address hallucinations caused by spurious correlations.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
Wang, Shaowen
Dong, Yiqi
Chang, Ruinian
Zhu, Tansheng
Sun, Yuebo
Lyu, Kaifeng
Li, Jian
Computation and Language
Artificial Intelligence
Machine Learning
Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and proprietary LLMs (including GPT-5), we show that existing hallucination detection methods, such as confidence-based filtering and inner-state probing, fundamentally fail in the presence of spurious correlations. Our theoretical analysis further elucidates why these statistical biases intrinsically undermine confidence-based detection techniques. Our findings thus emphasize the urgent need for new approaches explicitly designed to address hallucinations caused by spurious correlations.
title When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.07318