Towards Unification of Hallucination Detection and Fact Verification for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Weihang, Long, Jianming, Wang, Changyue, Lin, Shiyu, Xu, Jingyan, Ye, Ziyi, Ai, Qingyao, Liu, Yiqun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911298007072768
author Su, Weihang
Long, Jianming
Wang, Changyue
Lin, Shiyu
Xu, Jingyan
Ye, Ziyi
Ai, Qingyao
Liu, Yiqun
author_facet Su, Weihang
Long, Jianming
Wang, Changyue
Lin, Shiyu
Xu, Jingyan
Ye, Ziyi
Ai, Qingyao
Liu, Yiqun
contents Large Language Models (LLMs) frequently exhibit hallucinations, generating content that appears fluent and coherent but is factually incorrect. Such errors undermine trust and hinder their adoption in real-world applications. To address this challenge, two distinct research paradigms have emerged: model-centric Hallucination Detection (HD) and text-centric Fact Verification (FV). Despite sharing the same goal, these paradigms have evolved in isolation, using distinct assumptions, datasets, and evaluation protocols. This separation has created a research schism that hinders their collective progress. In this work, we take a decisive step toward bridging this divide. We introduce UniFact, a unified evaluation framework that enables direct, instance-level comparison between FV and HD by dynamically generating model outputs and corresponding factuality labels. Through large-scale experiments across multiple LLM families and detection methods, we reveal three key findings: (1) No paradigm is universally superior; (2) HD and FV capture complementary facets of factual errors; and (3) hybrid approaches that integrate both methods consistently achieve state-of-the-art performance. Beyond benchmarking, we provide the first in-depth analysis of why FV and HD diverged, as well as empirical evidence supporting the need for their unification. The comprehensive experimental results call for a new, integrated research agenda toward unifying Hallucination Detection and Fact Verification in LLMs. We have open-sourced all the code, data, and baseline implementation at: https://github.com/oneal2000/UniFact/
format Preprint
id arxiv_https___arxiv_org_abs_2512_02772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
Su, Weihang
Long, Jianming
Wang, Changyue
Lin, Shiyu
Xu, Jingyan
Ye, Ziyi
Ai, Qingyao
Liu, Yiqun
Computation and Language
Information Retrieval
Large Language Models (LLMs) frequently exhibit hallucinations, generating content that appears fluent and coherent but is factually incorrect. Such errors undermine trust and hinder their adoption in real-world applications. To address this challenge, two distinct research paradigms have emerged: model-centric Hallucination Detection (HD) and text-centric Fact Verification (FV). Despite sharing the same goal, these paradigms have evolved in isolation, using distinct assumptions, datasets, and evaluation protocols. This separation has created a research schism that hinders their collective progress. In this work, we take a decisive step toward bridging this divide. We introduce UniFact, a unified evaluation framework that enables direct, instance-level comparison between FV and HD by dynamically generating model outputs and corresponding factuality labels. Through large-scale experiments across multiple LLM families and detection methods, we reveal three key findings: (1) No paradigm is universally superior; (2) HD and FV capture complementary facets of factual errors; and (3) hybrid approaches that integrate both methods consistently achieve state-of-the-art performance. Beyond benchmarking, we provide the first in-depth analysis of why FV and HD diverged, as well as empirical evidence supporting the need for their unification. The comprehensive experimental results call for a new, integrated research agenda toward unifying Hallucination Detection and Fact Verification in LLMs. We have open-sourced all the code, data, and baseline implementation at: https://github.com/oneal2000/UniFact/
title Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2512.02772