Less is More for Long Document Summary Evaluation by LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Yunshu, Iso, Hayate, Pezeshkpour, Pouya, Bhutani, Nikita, Hruschka, Estevam
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916095812698112
author Wu, Yunshu
Iso, Hayate
Pezeshkpour, Pouya
Bhutani, Nikita
Hruschka, Estevam
author_facet Wu, Yunshu
Iso, Hayate
Pezeshkpour, Pouya
Bhutani, Nikita
Hruschka, Estevam
contents Large Language Models (LLMs) have shown promising performance in summary evaluation tasks, yet they face challenges such as high computational costs and the Lost-in-the-Middle problem where important information in the middle of long documents is often overlooked. To address these issues, this paper introduces a novel approach, Extract-then-Evaluate, which involves extracting key sentences from a long source document and then evaluating the summary by prompting LLMs. The results reveal that the proposed method not only significantly reduces evaluation costs but also exhibits a higher correlation with human evaluations. Furthermore, we provide practical recommendations for optimal document length and sentence extraction methods, contributing to the development of cost-effective yet more accurate methods for LLM-based text generation evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2309_07382
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Less is More for Long Document Summary Evaluation by LLMs
Wu, Yunshu
Iso, Hayate
Pezeshkpour, Pouya
Bhutani, Nikita
Hruschka, Estevam
Computation and Language
Large Language Models (LLMs) have shown promising performance in summary evaluation tasks, yet they face challenges such as high computational costs and the Lost-in-the-Middle problem where important information in the middle of long documents is often overlooked. To address these issues, this paper introduces a novel approach, Extract-then-Evaluate, which involves extracting key sentences from a long source document and then evaluating the summary by prompting LLMs. The results reveal that the proposed method not only significantly reduces evaluation costs but also exhibits a higher correlation with human evaluations. Furthermore, we provide practical recommendations for optimal document length and sentence extraction methods, contributing to the development of cost-effective yet more accurate methods for LLM-based text generation evaluation.
title Less is More for Long Document Summary Evaluation by LLMs
topic Computation and Language
url https://arxiv.org/abs/2309.07382