HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yeh, Min-Hsuan, Kamachee, Max, Park, Seongheon, Li, Yixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911137689239552
author Yeh, Min-Hsuan
Kamachee, Max
Park, Seongheon
Li, Yixuan
author_facet Yeh, Min-Hsuan
Kamachee, Max
Park, Seongheon
Li, Yixuan
contents To mitigate the impact of hallucination nature of LLMs, many studies propose detecting hallucinated generation through uncertainty estimation. However, these approaches predominantly operate at the sentence or paragraph level, failing to pinpoint specific spans or entities responsible for hallucinated content. This lack of granularity is especially problematic for long-form outputs that mix accurate and fabricated information. To address this limitation, we explore entity-level hallucination detection. We propose a new data set, HalluEntity, which annotates hallucination at the entity level. Based on the dataset, we comprehensively evaluate uncertainty-based hallucination detection approaches across 17 modern LLMs. Our experimental results show that uncertainty estimation approaches focusing on individual token probabilities tend to over-predict hallucinations, while context-aware methods show better but still suboptimal performance. Through an in-depth qualitative study, we identify relationships between hallucination tendencies and linguistic properties and highlight important directions for future research. HalluEntity: https://huggingface.co/datasets/samuelyeh/HalluEntity
format Preprint
id arxiv_https___arxiv_org_abs_2502_11948
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
Yeh, Min-Hsuan
Kamachee, Max
Park, Seongheon
Li, Yixuan
Computation and Language
To mitigate the impact of hallucination nature of LLMs, many studies propose detecting hallucinated generation through uncertainty estimation. However, these approaches predominantly operate at the sentence or paragraph level, failing to pinpoint specific spans or entities responsible for hallucinated content. This lack of granularity is especially problematic for long-form outputs that mix accurate and fabricated information. To address this limitation, we explore entity-level hallucination detection. We propose a new data set, HalluEntity, which annotates hallucination at the entity level. Based on the dataset, we comprehensively evaluate uncertainty-based hallucination detection approaches across 17 modern LLMs. Our experimental results show that uncertainty estimation approaches focusing on individual token probabilities tend to over-predict hallucinations, while context-aware methods show better but still suboptimal performance. Through an in-depth qualitative study, we identify relationships between hallucination tendencies and linguistic properties and highlight important directions for future research. HalluEntity: https://huggingface.co/datasets/samuelyeh/HalluEntity
title HalluEntity: Benchmarking and Understanding Entity-Level Hallucination Detection
topic Computation and Language
url https://arxiv.org/abs/2502.11948