RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pellegrini, Chantal, Özsoy, Ege, Busam, Benjamin, Navab, Nassir, Keicher, Matthias
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912363875139584
author Pellegrini, Chantal
Özsoy, Ege
Busam, Benjamin
Navab, Nassir
Keicher, Matthias
author_facet Pellegrini, Chantal
Özsoy, Ege
Busam, Benjamin
Navab, Nassir
Keicher, Matthias
contents Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-loop radiology assistant could facilitate a collaborative diagnostic process, thus saving time and improving the quality of reports. Towards this goal, we introduce RaDialog, the first thoroughly evaluated and publicly available large vision-language model for radiology report generation and interactive dialog. RaDialog effectively integrates visual image features and structured pathology findings with a large language model (LLM) while simultaneously adapting it to a specialized domain using parameter-efficient fine-tuning. To keep the conversational abilities of the underlying LLM, we propose a comprehensive, semi-automatically labeled, image-grounded instruct dataset for chest X-ray radiology tasks. By training with this dataset, our method achieves state-of-the-art clinical correctness in report generation and shows impressive abilities in interactive tasks such as correcting reports and answering questions, serving as a foundational step toward clinical dialog systems. Our code is available on github: https://github.com/ChantalMP/RaDialog.
format Preprint
id arxiv_https___arxiv_org_abs_2311_18681
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
Pellegrini, Chantal
Özsoy, Ege
Busam, Benjamin
Navab, Nassir
Keicher, Matthias
Computer Vision and Pattern Recognition
Computation and Language
Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-loop radiology assistant could facilitate a collaborative diagnostic process, thus saving time and improving the quality of reports. Towards this goal, we introduce RaDialog, the first thoroughly evaluated and publicly available large vision-language model for radiology report generation and interactive dialog. RaDialog effectively integrates visual image features and structured pathology findings with a large language model (LLM) while simultaneously adapting it to a specialized domain using parameter-efficient fine-tuning. To keep the conversational abilities of the underlying LLM, we propose a comprehensive, semi-automatically labeled, image-grounded instruct dataset for chest X-ray radiology tasks. By training with this dataset, our method achieves state-of-the-art clinical correctness in report generation and shows impressive abilities in interactive tasks such as correcting reports and answering questions, serving as a foundational step toward clinical dialog systems. Our code is available on github: https://github.com/ChantalMP/RaDialog.
title RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2311.18681