HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Fan, Hu, Xinyu, Yu, Zhenghan, Lin, Li, Zhang, Xu, Zhang, Yang, Zhou, Wei, Gu, Jinjie, Wan, Xiaojun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917032546533376
author Xu, Fan
Hu, Xinyu
Yu, Zhenghan
Lin, Li
Zhang, Xu
Zhang, Yang
Zhou, Wei
Gu, Jinjie
Wan, Xiaojun
author_facet Xu, Fan
Hu, Xinyu
Yu, Zhenghan
Lin, Li
Zhang, Xu
Zhang, Yang
Zhou, Wei
Gu, Jinjie
Wan, Xiaojun
contents The increasing reliance on natural language generation (NLG) models, particularly large language models, has raised concerns about the reliability and accuracy of their outputs. A key challenge is hallucination, where models produce plausible but incorrect information. As a result, hallucination detection has become a critical task. In this work, we introduce a comprehensive hallucination taxonomy with 11 categories across various NLG tasks and propose the HAllucination Detection (HAD) models https://github.com/pku0xff/HAD, which integrate hallucination detection, span-level identification, and correction into a single inference process. Trained on an elaborate synthetic dataset of about 90K samples, our HAD models are versatile and can be applied to various NLG tasks. We also carefully annotate a test set for hallucination detection, called HADTest, which contains 2,248 samples. Evaluations on in-domain and out-of-domain test sets show that our HAD models generally outperform the existing baselines, achieving state-of-the-art results on HaluEval, FactCHD, and FaithBench, confirming their robustness and versatility.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19318
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
Xu, Fan
Hu, Xinyu
Yu, Zhenghan
Lin, Li
Zhang, Xu
Zhang, Yang
Zhou, Wei
Gu, Jinjie
Wan, Xiaojun
Computation and Language
The increasing reliance on natural language generation (NLG) models, particularly large language models, has raised concerns about the reliability and accuracy of their outputs. A key challenge is hallucination, where models produce plausible but incorrect information. As a result, hallucination detection has become a critical task. In this work, we introduce a comprehensive hallucination taxonomy with 11 categories across various NLG tasks and propose the HAllucination Detection (HAD) models https://github.com/pku0xff/HAD, which integrate hallucination detection, span-level identification, and correction into a single inference process. Trained on an elaborate synthetic dataset of about 90K samples, our HAD models are versatile and can be applied to various NLG tasks. We also carefully annotate a test set for hallucination detection, called HADTest, which contains 2,248 samples. Evaluations on in-domain and out-of-domain test sets show that our HAD models generally outperform the existing baselines, achieving state-of-the-art results on HaluEval, FactCHD, and FaithBench, confirming their robustness and versatility.
title HAD: HAllucination Detection Language Models Based on a Comprehensive Hallucination Taxonomy
topic Computation and Language
url https://arxiv.org/abs/2510.19318