Localizing and Mitigating Errors in Long-form Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sachdeva, Rachneet, Song, Yixiao, Iyyer, Mohit, Gurevych, Iryna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910981198708736
author Sachdeva, Rachneet
Song, Yixiao
Iyyer, Mohit
Gurevych, Iryna
author_facet Sachdeva, Rachneet
Song, Yixiao
Iyyer, Mohit
Gurevych, Iryna
contents Long-form question answering (LFQA) aims to provide thorough and in-depth answers to complex questions, enhancing comprehension. However, such detailed responses are prone to hallucinations and factual inconsistencies, challenging their faithful evaluation. This work introduces HaluQuestQA, the first hallucination dataset with localized error annotations for human-written and model-generated LFQA answers. HaluQuestQA comprises 698 QA pairs with 1.8k span-level error annotations for five different error types by expert annotators, along with preference judgments. Using our collected data, we thoroughly analyze the shortcomings of long-form answers and find that they lack comprehensiveness and provide unhelpful references. We train an automatic feedback model on this dataset that predicts error spans with incomplete information and provides associated explanations. Finally, we propose a prompt-based approach, Error-informed refinement, that uses signals from the learned feedback model to refine generated answers, which we show reduces errors and improves answer quality across multiple models. Furthermore, humans find answers generated by our approach comprehensive and highly prefer them (84%) over the baseline answers.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Localizing and Mitigating Errors in Long-form Question Answering
Sachdeva, Rachneet
Song, Yixiao
Iyyer, Mohit
Gurevych, Iryna
Computation and Language
Long-form question answering (LFQA) aims to provide thorough and in-depth answers to complex questions, enhancing comprehension. However, such detailed responses are prone to hallucinations and factual inconsistencies, challenging their faithful evaluation. This work introduces HaluQuestQA, the first hallucination dataset with localized error annotations for human-written and model-generated LFQA answers. HaluQuestQA comprises 698 QA pairs with 1.8k span-level error annotations for five different error types by expert annotators, along with preference judgments. Using our collected data, we thoroughly analyze the shortcomings of long-form answers and find that they lack comprehensiveness and provide unhelpful references. We train an automatic feedback model on this dataset that predicts error spans with incomplete information and provides associated explanations. Finally, we propose a prompt-based approach, Error-informed refinement, that uses signals from the learned feedback model to refine generated answers, which we show reduces errors and improves answer quality across multiple models. Furthermore, humans find answers generated by our approach comprehensive and highly prefer them (84%) over the baseline answers.
title Localizing and Mitigating Errors in Long-form Question Answering
topic Computation and Language
url https://arxiv.org/abs/2407.11930