Fine-grained Hallucination Detection and Editing for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mishra, Abhika, Asai, Akari, Balachandran, Vidhisha, Wang, Yizhong, Neubig, Graham, Tsvetkov, Yulia, Hajishirzi, Hannaneh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914910582079488
author Mishra, Abhika
Asai, Akari
Balachandran, Vidhisha
Wang, Yizhong
Neubig, Graham
Tsvetkov, Yulia
Hajishirzi, Hannaneh
author_facet Mishra, Abhika
Asai, Akari
Balachandran, Vidhisha
Wang, Yizhong
Neubig, Graham
Tsvetkov, Yulia
Hajishirzi, Hannaneh
contents Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in diverse forms, each requiring varying degrees of careful assessments to verify factuality. We propose a novel task of automatic fine-grained hallucination detection and construct a new evaluation benchmark, FavaBench, that includes about one thousand fine-grained human judgments on three LM outputs across various domains. Our analysis reveals that ChatGPT and Llama2-Chat (70B, 7B) exhibit diverse types of hallucinations in the majority of their outputs in information-seeking scenarios. We train FAVA, a retrieval-augmented LM by carefully creating synthetic data to detect and correct fine-grained hallucinations. On our benchmark, our automatic and human evaluations show that FAVA significantly outperforms ChatGPT and GPT-4 on fine-grained hallucination detection, and edits suggested by FAVA improve the factuality of LM-generated text.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06855
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fine-grained Hallucination Detection and Editing for Language Models
Mishra, Abhika
Asai, Akari
Balachandran, Vidhisha
Wang, Yizhong
Neubig, Graham
Tsvetkov, Yulia
Hajishirzi, Hannaneh
Computation and Language
Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in diverse forms, each requiring varying degrees of careful assessments to verify factuality. We propose a novel task of automatic fine-grained hallucination detection and construct a new evaluation benchmark, FavaBench, that includes about one thousand fine-grained human judgments on three LM outputs across various domains. Our analysis reveals that ChatGPT and Llama2-Chat (70B, 7B) exhibit diverse types of hallucinations in the majority of their outputs in information-seeking scenarios. We train FAVA, a retrieval-augmented LM by carefully creating synthetic data to detect and correct fine-grained hallucinations. On our benchmark, our automatic and human evaluations show that FAVA significantly outperforms ChatGPT and GPT-4 on fine-grained hallucination detection, and edits suggested by FAVA improve the factuality of LM-generated text.
title Fine-grained Hallucination Detection and Editing for Language Models
topic Computation and Language
url https://arxiv.org/abs/2401.06855