InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trienes, Jan, Joseph, Sebastian, Schlötterer, Jörg, Seifert, Christin, Lo, Kyle, Xu, Wei, Wallace, Byron C., Li, Junyi Jessy
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911902900158464
author Trienes, Jan
Joseph, Sebastian
Schlötterer, Jörg
Seifert, Christin
Lo, Kyle
Xu, Wei
Wallace, Byron C.
Li, Junyi Jessy
author_facet Trienes, Jan
Joseph, Sebastian
Schlötterer, Jörg
Seifert, Christin
Lo, Kyle
Xu, Wei
Wallace, Byron C.
Li, Junyi Jessy
contents Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes InfoLossQA, a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs. Building on the theory of Question Under Discussion, the QA pairs are designed to help readers deepen their knowledge of a text. We conduct a range of experiments with this framework. First, we collect a dataset of 1,000 linguist-curated QA pairs derived from 104 LLM simplifications of scientific abstracts of medical studies. Our analyses of this data reveal that information loss occurs frequently, and that the QA pairs give a high-level overview of what information was lost. Second, we devise two methods for this task: end-to-end prompting of open-source and commercial language models, and a natural language inference pipeline. With a novel evaluation framework considering the correctness of QA pairs and their linguistic suitability, our expert evaluation reveals that models struggle to reliably identify information loss and applying similar standards as humans at what constitutes information loss.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16475
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
Trienes, Jan
Joseph, Sebastian
Schlötterer, Jörg
Seifert, Christin
Lo, Kyle
Xu, Wei
Wallace, Byron C.
Li, Junyi Jessy
Computation and Language
Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness. This work proposes InfoLossQA, a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs. Building on the theory of Question Under Discussion, the QA pairs are designed to help readers deepen their knowledge of a text. We conduct a range of experiments with this framework. First, we collect a dataset of 1,000 linguist-curated QA pairs derived from 104 LLM simplifications of scientific abstracts of medical studies. Our analyses of this data reveal that information loss occurs frequently, and that the QA pairs give a high-level overview of what information was lost. Second, we devise two methods for this task: end-to-end prompting of open-source and commercial language models, and a natural language inference pipeline. With a novel evaluation framework considering the correctness of QA pairs and their linguistic suitability, our expert evaluation reveals that models struggle to reliably identify information loss and applying similar standards as humans at what constitutes information loss.
title InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
topic Computation and Language
url https://arxiv.org/abs/2401.16475