Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Samarinas, Chris, Krubner, Alexander, Salemi, Alireza, Kim, Youngwoo, Zamani, Hamed
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916770035531776
author Samarinas, Chris
Krubner, Alexander
Salemi, Alireza
Kim, Youngwoo
Zamani, Hamed
author_facet Samarinas, Chris
Krubner, Alexander
Salemi, Alireza
Kim, Youngwoo
Zamani, Hamed
contents This paper presents ICAT, an evaluation framework for measuring coverage of diverse factual information in long-form text generation. ICAT breaks down a long output text into a list of atomic claims and not only verifies each claim through retrieval from a (reliable) knowledge source, but also computes the alignment between the atomic factual claims and various aspects expected to be presented in the output. We study three implementations of the ICAT framework, each with a different assumption on the availability of aspects and alignment method. By adopting data from the diversification task in the TREC Web Track and the ClueWeb corpus, we evaluate the ICAT framework. We demonstrate strong correlation with human judgments and provide comprehensive evaluation across multiple state-of-the-art LLMs. Our framework further offers interpretable and fine-grained analysis of diversity and coverage. Its modular design allows for easy adaptation to different domains and datasets, making it a valuable tool for evaluating the qualitative aspects of long-form responses produced by LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2501_03545
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
Samarinas, Chris
Krubner, Alexander
Salemi, Alireza
Kim, Youngwoo
Zamani, Hamed
Computation and Language
This paper presents ICAT, an evaluation framework for measuring coverage of diverse factual information in long-form text generation. ICAT breaks down a long output text into a list of atomic claims and not only verifies each claim through retrieval from a (reliable) knowledge source, but also computes the alignment between the atomic factual claims and various aspects expected to be presented in the output. We study three implementations of the ICAT framework, each with a different assumption on the availability of aspects and alignment method. By adopting data from the diversification task in the TREC Web Track and the ClueWeb corpus, we evaluate the ICAT framework. We demonstrate strong correlation with human judgments and provide comprehensive evaluation across multiple state-of-the-art LLMs. Our framework further offers interpretable and fine-grained analysis of diversity and coverage. Its modular design allows for easy adaptation to different domains and datasets, making it a valuable tool for evaluating the qualitative aspects of long-form responses produced by LLMs.
title Beyond Factual Accuracy: Evaluating Coverage of Diverse Factual Information in Long-form Text Generation
topic Computation and Language
url https://arxiv.org/abs/2501.03545