LCFO: Long Context and Long Form Output Dataset and Benchmarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Costa-jussà, Marta R., Andrews, Pierre, Meglioli, Mariano Coria, Chen, Joy, Chuang, Joe, Dale, David, Ropers, Christophe, Mourachko, Alexandre, Sánchez, Eduardo, Schwenk, Holger, Tran, Tuan, Turkatenko, Arina, Wood, Carleigh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911046205177856
author Costa-jussà, Marta R.
Andrews, Pierre
Meglioli, Mariano Coria
Chen, Joy
Chuang, Joe
Dale, David
Ropers, Christophe
Mourachko, Alexandre
Sánchez, Eduardo
Schwenk, Holger
Tran, Tuan
Turkatenko, Arina
Wood, Carleigh
author_facet Costa-jussà, Marta R.
Andrews, Pierre
Meglioli, Mariano Coria
Chen, Joy
Chuang, Joe
Dale, David
Ropers, Christophe
Mourachko, Alexandre
Sánchez, Eduardo
Schwenk, Holger
Tran, Tuan
Turkatenko, Arina
Wood, Carleigh
contents This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries of different lengths (20%, 10%, and 5% of the input text), as well as approximately 15 questions and answers (QA) related to the input content. Notably, LCFO also provides alignments between specific QA pairs and corresponding summaries in 7 domains. The primary motivation behind providing summaries of different lengths is to establish a controllable framework for generating long texts from shorter inputs, i.e. summary expansion. To establish an evaluation metric framework for summarization and summary expansion, we provide human evaluation scores for human-generated outputs, as well as results from various state-of-the-art large language models (LLMs). GPT-4o-mini achieves best human scores among automatic systems in both summarization and summary expansion tasks (~ +10% and +20%, respectively). It even surpasses human output quality in the case of short summaries (~ +7%). Overall automatic metrics achieve low correlations with human evaluation scores (~ 0.4) but moderate correlation on specific evaluation aspects such as fluency and attribution (~ 0.6).
format Preprint
id arxiv_https___arxiv_org_abs_2412_08268
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LCFO: Long Context and Long Form Output Dataset and Benchmarking
Costa-jussà, Marta R.
Andrews, Pierre
Meglioli, Mariano Coria
Chen, Joy
Chuang, Joe
Dale, David
Ropers, Christophe
Mourachko, Alexandre
Sánchez, Eduardo
Schwenk, Holger
Tran, Tuan
Turkatenko, Arina
Wood, Carleigh
Computation and Language
I.2.7
This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries of different lengths (20%, 10%, and 5% of the input text), as well as approximately 15 questions and answers (QA) related to the input content. Notably, LCFO also provides alignments between specific QA pairs and corresponding summaries in 7 domains. The primary motivation behind providing summaries of different lengths is to establish a controllable framework for generating long texts from shorter inputs, i.e. summary expansion. To establish an evaluation metric framework for summarization and summary expansion, we provide human evaluation scores for human-generated outputs, as well as results from various state-of-the-art large language models (LLMs). GPT-4o-mini achieves best human scores among automatic systems in both summarization and summary expansion tasks (~ +10% and +20%, respectively). It even surpasses human output quality in the case of short summaries (~ +7%). Overall automatic metrics achieve low correlations with human evaluation scores (~ 0.4) but moderate correlation on specific evaluation aspects such as fluency and attribution (~ 0.6).
title LCFO: Long Context and Long Form Output Dataset and Benchmarking
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2412.08268