Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asthana, Sumit, Rashkin, Hannah, Clark, Elizabeth, Huot, Fantine, Lapata, Mirella
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912203181916160
author Asthana, Sumit
Rashkin, Hannah
Clark, Elizabeth
Huot, Fantine
Lapata, Mirella
author_facet Asthana, Sumit
Rashkin, Hannah
Clark, Elizabeth
Huot, Fantine
Lapata, Mirella
contents One useful application of NLP models is to support people in reading complex text from unfamiliar domains (e.g., scientific articles). Simplifying the entire text makes it understandable but sometimes removes important details. On the contrary, helping adult readers understand difficult concepts in context can enhance their vocabulary and knowledge. In a preliminary human study, we first identify that lack of context and unfamiliarity with difficult concepts is a major reason for adult readers' difficulty with domain-specific text. We then introduce "targeted concept simplification," a simplification task for rewriting text to help readers comprehend text containing unfamiliar concepts. We also introduce WikiDomains, a new dataset of 22k definitions from 13 academic domains paired with a difficult concept within each definition. We benchmark the performance of open-source and commercial LLMs and a simple dictionary baseline on this task across human judgments of ease of understanding and meaning preservation. Interestingly, our human judges preferred explanations about the difficult concept more than simplification of the concept phrase. Further, no single model achieved superior performance across all quality dimensions, and automated metrics also show low correlations with human evaluations of concept simplification ($\sim0.2$), opening up rich avenues for research on personalized human reading comprehension support.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
Asthana, Sumit
Rashkin, Hannah
Clark, Elizabeth
Huot, Fantine
Lapata, Mirella
Computation and Language
One useful application of NLP models is to support people in reading complex text from unfamiliar domains (e.g., scientific articles). Simplifying the entire text makes it understandable but sometimes removes important details. On the contrary, helping adult readers understand difficult concepts in context can enhance their vocabulary and knowledge. In a preliminary human study, we first identify that lack of context and unfamiliarity with difficult concepts is a major reason for adult readers' difficulty with domain-specific text. We then introduce "targeted concept simplification," a simplification task for rewriting text to help readers comprehend text containing unfamiliar concepts. We also introduce WikiDomains, a new dataset of 22k definitions from 13 academic domains paired with a difficult concept within each definition. We benchmark the performance of open-source and commercial LLMs and a simple dictionary baseline on this task across human judgments of ease of understanding and meaning preservation. Interestingly, our human judges preferred explanations about the difficult concept more than simplification of the concept phrase. Further, no single model achieved superior performance across all quality dimensions, and automated metrics also show low correlations with human evaluations of concept simplification ($\sim0.2$), opening up rich avenues for research on personalized human reading comprehension support.
title Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts
topic Computation and Language
url https://arxiv.org/abs/2410.20763