Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Zhenliang, Zhang, Junzhe, Hu, Xinyu, Zhang, HuiXuan, Wan, Xiaojun
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911101280583680
author Zhang, Zhenliang
Zhang, Junzhe
Hu, Xinyu
Zhang, HuiXuan
Wan, Xiaojun
author_facet Zhang, Zhenliang
Zhang, Junzhe
Hu, Xinyu
Zhang, HuiXuan
Wan, Xiaojun
contents Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the input. In this study, we investigate whether social bias contributes to these hallucinations, a causal relationship that has not been explored. A key challenge is controlling confounders within the context, which complicates the isolation of causality between bias states and hallucinations. To address this, we utilize the Structural Causal Model (SCM) to establish and validate the causality and design bias interventions to control confounders. In addition, we develop the Bias Intervention Dataset (BID), which includes various social biases, enabling precise measurement of causal effects. Experiments on mainstream LLMs reveal that biases are significant causes of faithfulness hallucinations, and the effect of each bias state differs in direction. We further analyze the scope of these causal effects across various models, specifically focusing on unfairness hallucinations, which are primarily targeted by social bias, revealing the subtle yet significant causal effect of bias on hallucination generation.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
Zhang, Zhenliang
Zhang, Junzhe
Hu, Xinyu
Zhang, HuiXuan
Wan, Xiaojun
Computation and Language
Large language models (LLMs) have achieved remarkable success in various tasks, yet they remain vulnerable to faithfulness hallucinations, where the output does not align with the input. In this study, we investigate whether social bias contributes to these hallucinations, a causal relationship that has not been explored. A key challenge is controlling confounders within the context, which complicates the isolation of causality between bias states and hallucinations. To address this, we utilize the Structural Causal Model (SCM) to establish and validate the causality and design bias interventions to control confounders. In addition, we develop the Bias Intervention Dataset (BID), which includes various social biases, enabling precise measurement of causal effects. Experiments on mainstream LLMs reveal that biases are significant causes of faithfulness hallucinations, and the effect of each bias state differs in direction. We further analyze the scope of these causal effects across various models, specifically focusing on unfairness hallucinations, which are primarily targeted by social bias, revealing the subtle yet significant causal effect of bias on hallucination generation.
title Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2508.07753