The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuji, Li, Sha, Qian, Cheng, Liu, Jiateng, Yu, Pengfei, Han, Chi, Fung, Yi R., McKeown, Kathleen, Zhai, Chengxiang, Li, Manling, Ji, Heng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912241939382272
author Zhang, Yuji
Li, Sha
Qian, Cheng
Liu, Jiateng
Yu, Pengfei
Han, Chi
Fung, Yi R.
McKeown, Kathleen
Zhai, Chengxiang
Li, Manling
Ji, Heng
author_facet Zhang, Yuji
Li, Sha
Qian, Cheng
Liu, Jiateng
Yu, Pengfei
Han, Chi
Fung, Yi R.
McKeown, Kathleen
Zhai, Chengxiang
Li, Manling
Ji, Heng
contents Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which error generation continues despite high-quality training data, calls for a deeper understanding of the underlying LLM mechanisms. To address it, we propose a novel concept: knowledge overshadowing, where model's dominant knowledge can obscure less prominent knowledge during text generation, causing the model to fabricate inaccurate details. Building on this idea, we introduce a novel framework to quantify factual hallucinations by modeling knowledge overshadowing. Central to our approach is the log-linear law, which predicts that the rate of factual hallucination increases linearly with the logarithmic scale of (1) Knowledge Popularity, (2) Knowledge Length, and (3) Model Size. The law provides a means to preemptively quantify hallucinations, offering foresight into their occurrence even before model training or inference. Built on overshadowing effect, we propose a new decoding strategy CoDa, to mitigate hallucinations, which notably enhance model factuality on Overshadow (27.9%), MemoTrap (13.1%) and NQ-Swap (18.3%). Our findings not only deepen understandings of the underlying mechanisms behind hallucinations but also provide actionable insights for developing more predictable and controllable language models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16143
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
Zhang, Yuji
Li, Sha
Qian, Cheng
Liu, Jiateng
Yu, Pengfei
Han, Chi
Fung, Yi R.
McKeown, Kathleen
Zhai, Chengxiang
Li, Manling
Ji, Heng
Computation and Language
Hallucination is a persistent challenge in large language models (LLMs), where even with rigorous quality control, models often generate distorted facts. This paradox, in which error generation continues despite high-quality training data, calls for a deeper understanding of the underlying LLM mechanisms. To address it, we propose a novel concept: knowledge overshadowing, where model's dominant knowledge can obscure less prominent knowledge during text generation, causing the model to fabricate inaccurate details. Building on this idea, we introduce a novel framework to quantify factual hallucinations by modeling knowledge overshadowing. Central to our approach is the log-linear law, which predicts that the rate of factual hallucination increases linearly with the logarithmic scale of (1) Knowledge Popularity, (2) Knowledge Length, and (3) Model Size. The law provides a means to preemptively quantify hallucinations, offering foresight into their occurrence even before model training or inference. Built on overshadowing effect, we propose a new decoding strategy CoDa, to mitigate hallucinations, which notably enhance model factuality on Overshadow (27.9%), MemoTrap (13.1%) and NQ-Swap (18.3%). Our findings not only deepen understandings of the underlying mechanisms behind hallucinations but also provide actionable insights for developing more predictable and controllable language models.
title The Law of Knowledge Overshadowing: Towards Understanding, Predicting, and Preventing LLM Hallucination
topic Computation and Language
url https://arxiv.org/abs/2502.16143