Real-World Summarization: When Evaluation Reaches Its Limits

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmidtová, Patrícia, Dušek, Ondřej, Mahamood, Saad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909690142654464
author Schmidtová, Patrícia
Dušek, Ondřej
Mahamood, Saad
author_facet Schmidtová, Patrícia
Dušek, Ondřej
Mahamood, Saad
contents We examine evaluation of faithfulness to input data in the context of hotel highlights: brief LLM-generated summaries that capture unique features of accommodations. Through human evaluation campaigns involving categorical error assessment and span-level annotation, we compare traditional metrics, trainable methods, and LLM-as-a-judge approaches. Our findings reveal that simpler metrics like word overlap correlate surprisingly well with human judgments (Spearman correlation rank of 0.63), often outperforming more complex methods when applied to out-of-domain data. We further demonstrate that while LLMs can generate high-quality highlights, they prove unreliable for evaluation as they tend to severely under- or over-annotate. Our analysis of real-world business impacts shows incorrect and non-checkable information pose the greatest risks. We also highlight challenges in crowdsourced evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2507_11508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Real-World Summarization: When Evaluation Reaches Its Limits
Schmidtová, Patrícia
Dušek, Ondřej
Mahamood, Saad
Computation and Language
We examine evaluation of faithfulness to input data in the context of hotel highlights: brief LLM-generated summaries that capture unique features of accommodations. Through human evaluation campaigns involving categorical error assessment and span-level annotation, we compare traditional metrics, trainable methods, and LLM-as-a-judge approaches. Our findings reveal that simpler metrics like word overlap correlate surprisingly well with human judgments (Spearman correlation rank of 0.63), often outperforming more complex methods when applied to out-of-domain data. We further demonstrate that while LLMs can generate high-quality highlights, they prove unreliable for evaluation as they tend to severely under- or over-annotate. Our analysis of real-world business impacts shows incorrect and non-checkable information pose the greatest risks. We also highlight challenges in crowdsourced evaluations.
title Real-World Summarization: When Evaluation Reaches Its Limits
topic Computation and Language
url https://arxiv.org/abs/2507.11508