SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Thorat, Onkar, Laban, Philippe, Wu, Chien-Sheng
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912404715077632
author Thorat, Onkar
Laban, Philippe
Wu, Chien-Sheng
author_facet Thorat, Onkar
Laban, Philippe
Wu, Chien-Sheng
contents Detecting factual inconsistencies in summarization is critical, yet existing benchmarks lack the necessary challenge and interpretability for robust evaluation. In this paper, we introduce SummExecEdit, a novel pipeline and benchmark leveraging executable edits to assess models on their ability to both detect factual errors and provide accurate explanations. The top-performing model, Claude3-Opus, achieves a joint detection and explanation score of only 0.49 in our benchmark, with individual scores of 0.67 for detection and 0.73 for explanation. We conduct detailed evaluations to assess the current state of models in this field and find that more than half of the 20+ LLMs in our study struggle with over 30% of the SummExecEdit benchmark. Additionally, we identify four primary types of explanation errors, with 45.4% of them involving a focus on completely unrelated parts of the summary.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13378
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
Thorat, Onkar
Laban, Philippe
Wu, Chien-Sheng
Computation and Language
Detecting factual inconsistencies in summarization is critical, yet existing benchmarks lack the necessary challenge and interpretability for robust evaluation. In this paper, we introduce SummExecEdit, a novel pipeline and benchmark leveraging executable edits to assess models on their ability to both detect factual errors and provide accurate explanations. The top-performing model, Claude3-Opus, achieves a joint detection and explanation score of only 0.49 in our benchmark, with individual scores of 0.67 for detection and 0.73 for explanation. We conduct detailed evaluations to assess the current state of models in this field and find that more than half of the 20+ LLMs in our study struggle with over 30% of the SummExecEdit benchmark. Additionally, we identify four primary types of explanation errors, with 45.4% of them involving a focus on completely unrelated parts of the summary.
title SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
topic Computation and Language
url https://arxiv.org/abs/2412.13378