Fine-grained Hallucination Detection and Editing for Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914910582079488 |
|---|---|
| author | Mishra, Abhika Asai, Akari Balachandran, Vidhisha Wang, Yizhong Neubig, Graham Tsvetkov, Yulia Hajishirzi, Hannaneh |
| author_facet | Mishra, Abhika Asai, Akari Balachandran, Vidhisha Wang, Yizhong Neubig, Graham Tsvetkov, Yulia Hajishirzi, Hannaneh |
| contents | Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in diverse forms, each requiring varying degrees of careful assessments to verify factuality. We propose a novel task of automatic fine-grained hallucination detection and construct a new evaluation benchmark, FavaBench, that includes about one thousand fine-grained human judgments on three LM outputs across various domains. Our analysis reveals that ChatGPT and Llama2-Chat (70B, 7B) exhibit diverse types of hallucinations in the majority of their outputs in information-seeking scenarios. We train FAVA, a retrieval-augmented LM by carefully creating synthetic data to detect and correct fine-grained hallucinations. On our benchmark, our automatic and human evaluations show that FAVA significantly outperforms ChatGPT and GPT-4 on fine-grained hallucination detection, and edits suggested by FAVA improve the factuality of LM-generated text. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2401_06855 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Fine-grained Hallucination Detection and Editing for Language Models Mishra, Abhika Asai, Akari Balachandran, Vidhisha Wang, Yizhong Neubig, Graham Tsvetkov, Yulia Hajishirzi, Hannaneh Computation and Language Large language models (LMs) are prone to generate factual errors, which are often called hallucinations. In this paper, we introduce a comprehensive taxonomy of hallucinations and argue that hallucinations manifest in diverse forms, each requiring varying degrees of careful assessments to verify factuality. We propose a novel task of automatic fine-grained hallucination detection and construct a new evaluation benchmark, FavaBench, that includes about one thousand fine-grained human judgments on three LM outputs across various domains. Our analysis reveals that ChatGPT and Llama2-Chat (70B, 7B) exhibit diverse types of hallucinations in the majority of their outputs in information-seeking scenarios. We train FAVA, a retrieval-augmented LM by carefully creating synthetic data to detect and correct fine-grained hallucinations. On our benchmark, our automatic and human evaluations show that FAVA significantly outperforms ChatGPT and GPT-4 on fine-grained hallucination detection, and edits suggested by FAVA improve the factuality of LM-generated text. |
| title | Fine-grained Hallucination Detection and Editing for Language Models |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2401.06855 |