ANAH: Analytical Annotation of Hallucinations in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Ziwei, Gu, Yuzhe, Zhang, Wenwei, Lyu, Chengqi, Lin, Dahua, Chen, Kai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910464546439168
author Ji, Ziwei
Gu, Yuzhe
Zhang, Wenwei
Lyu, Chengqi
Lin, Dahua
Chen, Kai
author_facet Ji, Ziwei
Gu, Yuzhe
Zhang, Wenwei
Lyu, Chengqi
Lin, Dahua
Chen, Kai
contents Reducing the `$\textit{hallucination}$' problem of Large Language Models (LLMs) is crucial for their wide applications. A comprehensive and fine-grained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community. Thus, we present $\textbf{ANAH}$, a bilingual dataset that offers $\textbf{AN}$alytical $\textbf{A}$nnotation of $\textbf{H}$allucinations in LLMs within Generative Question Answering. Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content. ANAH consists of ~12k sentence-level annotations for ~4.3k LLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline. Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators. We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ANAH: Analytical Annotation of Hallucinations in Large Language Models
Ji, Ziwei
Gu, Yuzhe
Zhang, Wenwei
Lyu, Chengqi
Lin, Dahua
Chen, Kai
Computation and Language
Artificial Intelligence
Reducing the `$\textit{hallucination}$' problem of Large Language Models (LLMs) is crucial for their wide applications. A comprehensive and fine-grained measurement of the hallucination is the first key step for the governance of this issue but is under-explored in the community. Thus, we present $\textbf{ANAH}$, a bilingual dataset that offers $\textbf{AN}$alytical $\textbf{A}$nnotation of $\textbf{H}$allucinations in LLMs within Generative Question Answering. Each answer sentence in our dataset undergoes rigorous annotation, involving the retrieval of a reference fragment, the judgment of the hallucination type, and the correction of hallucinated content. ANAH consists of ~12k sentence-level annotations for ~4.3k LLM responses covering over 700 topics, constructed by a human-in-the-loop pipeline. Thanks to the fine granularity of the hallucination annotations, we can quantitatively confirm that the hallucinations of LLMs progressively accumulate in the answer and use ANAH to train and evaluate hallucination annotators. We conduct extensive experiments on studying generative and discriminative annotators and show that, although current open-source LLMs have difficulties in fine-grained hallucination annotation, the generative annotator trained with ANAH can surpass all open-source LLMs and GPT-3.5, obtain performance competitive with GPT-4, and exhibits better generalization ability on unseen questions.
title ANAH: Analytical Annotation of Hallucinations in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2405.20315