Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Subbiah, Melanie, Mishra, Akankshya, Kim, Grace, Tang, Liyan, Durrett, Greg, McKeown, Kathleen
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911156919074816
author Subbiah, Melanie
Mishra, Akankshya
Kim, Grace
Tang, Liyan
Durrett, Greg
McKeown, Kathleen
author_facet Subbiah, Melanie
Mishra, Akankshya
Kim, Grace
Tang, Liyan
Durrett, Greg
McKeown, Kathleen
contents Determining faithfulness of a claim to a source document is an important problem across many domains. This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source. In many cases, though, whether a claim is supported can be ambiguous. For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences. Forcing binary labels upon such claims lowers the reliability of evaluation. In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims. We introduce LLM-generated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous? Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness. We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation. We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.
format Preprint
id arxiv_https___arxiv_org_abs_2504_01132
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
Subbiah, Melanie
Mishra, Akankshya
Kim, Grace
Tang, Liyan
Durrett, Greg
McKeown, Kathleen
Computation and Language
Artificial Intelligence
Determining faithfulness of a claim to a source document is an important problem across many domains. This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source. In many cases, though, whether a claim is supported can be ambiguous. For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences. Forcing binary labels upon such claims lowers the reliability of evaluation. In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims. We introduce LLM-generated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous? Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness. We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation. We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.
title Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.01132