Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Chenyu, Zhou, Weichao, Ghosh, Shantanu, Batmanghelich, Kayhan, Li, Wenchao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915199485739008
author Wang, Chenyu
Zhou, Weichao
Ghosh, Shantanu
Batmanghelich, Kayhan
Li, Wenchao
author_facet Wang, Chenyu
Zhou, Weichao
Ghosh, Shantanu
Batmanghelich, Kayhan
Li, Wenchao
contents Radiology report generation (RRG) has shown great potential in assisting radiologists by automating the labor-intensive task of report writing. While recent advancements have improved the quality and coherence of generated reports, ensuring their factual correctness remains a critical challenge. Although generative medical Vision Large Language Models (VLLMs) have been proposed to address this issue, these models are prone to hallucinations and can produce inaccurate diagnostic information. To address these concerns, we introduce a novel Semantic Consistency-Based Uncertainty Quantification framework that provides both report-level and sentence-level uncertainties. Unlike existing approaches, our method does not require modifications to the underlying model or access to its inner state, such as output token logits, thus serving as a plug-and-play module that can be seamlessly integrated with state-of-the-art models. Extensive experiments demonstrate the efficacy of our method in detecting hallucinations and enhancing the factual accuracy of automatically generated radiology reports. By abstaining from high-uncertainty reports, our approach improves factuality scores by $10$\%, achieved by rejecting $20$\% of reports using the \texttt{Radialog} model on the MIMIC-CXR dataset. Furthermore, sentence-level uncertainty flags the lowest-precision sentence in each report with an $82.9$\% success rate. Our implementation is open-source and available at https://github.com/BU-DEPEND-Lab/SCUQ-RRG.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04606
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
Wang, Chenyu
Zhou, Weichao
Ghosh, Shantanu
Batmanghelich, Kayhan
Li, Wenchao
Artificial Intelligence
Computation and Language
Radiology report generation (RRG) has shown great potential in assisting radiologists by automating the labor-intensive task of report writing. While recent advancements have improved the quality and coherence of generated reports, ensuring their factual correctness remains a critical challenge. Although generative medical Vision Large Language Models (VLLMs) have been proposed to address this issue, these models are prone to hallucinations and can produce inaccurate diagnostic information. To address these concerns, we introduce a novel Semantic Consistency-Based Uncertainty Quantification framework that provides both report-level and sentence-level uncertainties. Unlike existing approaches, our method does not require modifications to the underlying model or access to its inner state, such as output token logits, thus serving as a plug-and-play module that can be seamlessly integrated with state-of-the-art models. Extensive experiments demonstrate the efficacy of our method in detecting hallucinations and enhancing the factual accuracy of automatically generated radiology reports. By abstaining from high-uncertainty reports, our approach improves factuality scores by $10$\%, achieved by rejecting $20$\% of reports using the \texttt{Radialog} model on the MIMIC-CXR dataset. Furthermore, sentence-level uncertainty flags the lowest-precision sentence in each report with an $82.9$\% success rate. Our implementation is open-source and available at https://github.com/BU-DEPEND-Lab/SCUQ-RRG.
title Semantic Consistency-Based Uncertainty Quantification for Factuality in Radiology Report Generation
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2412.04606