Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gero, Zelalem, Singh, Chandan, Xie, Yiqing, Zhang, Sheng, Subramanian, Praveen, Vozila, Paul, Naumann, Tristan, Gao, Jianfeng, Poon, Hoifung
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915063110041600
author Gero, Zelalem
Singh, Chandan
Xie, Yiqing
Zhang, Sheng
Subramanian, Praveen
Vozila, Paul
Naumann, Tristan
Gao, Jianfeng
Poon, Hoifung
author_facet Gero, Zelalem
Singh, Chandan
Xie, Yiqing
Zhang, Sheng
Subramanian, Praveen
Vozila, Paul
Naumann, Tristan
Gao, Jianfeng
Poon, Hoifung
contents Summarizing clinical text is crucial in health decision-support and clinical research. Large language models (LLMs) have shown the potential to generate accurate clinical text summaries, but still struggle with issues regarding grounding and evaluation, especially in safety-critical domains such as health. Holistically evaluating text summaries is challenging because they may contain unsubstantiated information. Here, we explore a general mitigation framework using Attribute Structuring (AS), which structures the summary evaluation process. It decomposes the evaluation process into a grounded procedure that uses an LLM for relatively simple structuring and scoring tasks, rather than the full task of holistic summary evaluation. Experiments show that AS consistently improves the correspondence between human annotations and automated metrics in clinical text summarization. Additionally, AS yields interpretations in the form of a short text span corresponding to each output, which enables efficient human auditing, paving the way towards trustworthy evaluation of clinical information in resource-constrained scenarios. We release our code, prompts, and an open-source benchmark at https://github.com/microsoft/attribute-structuring.
format Preprint
id arxiv_https___arxiv_org_abs_2403_01002
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
Gero, Zelalem
Singh, Chandan
Xie, Yiqing
Zhang, Sheng
Subramanian, Praveen
Vozila, Paul
Naumann, Tristan
Gao, Jianfeng
Poon, Hoifung
Computation and Language
Artificial Intelligence
Summarizing clinical text is crucial in health decision-support and clinical research. Large language models (LLMs) have shown the potential to generate accurate clinical text summaries, but still struggle with issues regarding grounding and evaluation, especially in safety-critical domains such as health. Holistically evaluating text summaries is challenging because they may contain unsubstantiated information. Here, we explore a general mitigation framework using Attribute Structuring (AS), which structures the summary evaluation process. It decomposes the evaluation process into a grounded procedure that uses an LLM for relatively simple structuring and scoring tasks, rather than the full task of holistic summary evaluation. Experiments show that AS consistently improves the correspondence between human annotations and automated metrics in clinical text summarization. Additionally, AS yields interpretations in the form of a short text span corresponding to each output, which enables efficient human auditing, paving the way towards trustworthy evaluation of clinical information in resource-constrained scenarios. We release our code, prompts, and an open-source benchmark at https://github.com/microsoft/attribute-structuring.
title Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.01002