A Unified Definition of Hallucination: It's The World Model, Stupid!

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Emmy, Gangal, Varun, Zou, Chelsea, Yu, Michael, Huang, Xiaoqi, Chang, Alex, Tao, Zhuofu, Singh, Karan, Kumar, Sachin, Feng, Steven Y.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908807140999168
author Liu, Emmy
Gangal, Varun
Zou, Chelsea
Yu, Michael
Huang, Xiaoqi
Chang, Alex
Tao, Zhuofu
Singh, Karan
Kumar, Sachin
Feng, Steven Y.
author_facet Liu, Emmy
Gangal, Varun
Zou, Chelsea
Yu, Michael
Huang, Xiaoqi
Chang, Alex
Tao, Zhuofu
Singh, Karan
Kumar, Sachin
Feng, Steven Y.
contents Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition wherein prior definitions are subsumed. We argue that hallucination can be unified by defining it as simply inaccurate (internal) world modeling, in a form where it is observable to the user. For example, stating a fact which contradicts a knowledge base OR producing a summary which contradicts the source. By varying the reference world model and conflict policy, our framework unifies prior definitions. We argue that this unified view is useful because it forces evaluations to clarify their assumed reference "world", distinguishes true hallucinations from planning or reward errors, and provides a common language for comparison across benchmarks and discussion of mitigation strategies. Building on this definition, we outline plans for a family of benchmarks using synthetic, fully specified reference world models to stress-test and improve world modeling components.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21577
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Unified Definition of Hallucination: It's The World Model, Stupid!
Liu, Emmy
Gangal, Varun
Zou, Chelsea
Yu, Michael
Huang, Xiaoqi
Chang, Alex
Tao, Zhuofu
Singh, Karan
Kumar, Sachin
Feng, Steven Y.
Computation and Language
Artificial Intelligence
Machine Learning
Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition wherein prior definitions are subsumed. We argue that hallucination can be unified by defining it as simply inaccurate (internal) world modeling, in a form where it is observable to the user. For example, stating a fact which contradicts a knowledge base OR producing a summary which contradicts the source. By varying the reference world model and conflict policy, our framework unifies prior definitions. We argue that this unified view is useful because it forces evaluations to clarify their assumed reference "world", distinguishes true hallucinations from planning or reward errors, and provides a common language for comparison across benchmarks and discussion of mitigation strategies. Building on this definition, we outline plans for a family of benchmarks using synthetic, fully specified reference world models to stress-test and improve world modeling components.
title A Unified Definition of Hallucination: It's The World Model, Stupid!
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.21577