MAIRA-1: A specialised large multimodal model for radiology report generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hyland, Stephanie L., Bannur, Shruthi, Bouzid, Kenza, Castro, Daniel C., Ranjit, Mercy, Schwaighofer, Anton, Pérez-García, Fernando, Salvatelli, Valentina, Srivastav, Shaury, Thieme, Anja, Codella, Noel, Lungren, Matthew P., Wetscherek, Maria Teodora, Oktay, Ozan, Alvarez-Valle, Javier
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909183075418112
author Hyland, Stephanie L.
Bannur, Shruthi
Bouzid, Kenza
Castro, Daniel C.
Ranjit, Mercy
Schwaighofer, Anton
Pérez-García, Fernando
Salvatelli, Valentina
Srivastav, Shaury
Thieme, Anja
Codella, Noel
Lungren, Matthew P.
Wetscherek, Maria Teodora
Oktay, Ozan
Alvarez-Valle, Javier
author_facet Hyland, Stephanie L.
Bannur, Shruthi
Bouzid, Kenza
Castro, Daniel C.
Ranjit, Mercy
Schwaighofer, Anton
Pérez-García, Fernando
Salvatelli, Valentina
Srivastav, Shaury
Thieme, Anja
Codella, Noel
Lungren, Matthew P.
Wetscherek, Maria Teodora
Oktay, Ozan
Alvarez-Valle, Javier
contents We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s) can be equipped with multimodal capabilities through alignment with pre-trained vision encoders. On natural images, this has been shown to allow multimodal models to gain image understanding and description capabilities. Our proposed model (MAIRA-1) leverages a CXR-specific image encoder in conjunction with a fine-tuned large language model based on Vicuna-7B, and text-based data augmentation, to produce reports with state-of-the-art quality. In particular, MAIRA-1 significantly improves on the radiologist-aligned RadCliQ metric and across all lexical metrics considered. Manual review of model outputs demonstrates promising fluency and accuracy of generated reports while uncovering failure modes not captured by existing evaluation practices. More information and resources can be found on the project website: https://aka.ms/maira.
format Preprint
id arxiv_https___arxiv_org_abs_2311_13668
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MAIRA-1: A specialised large multimodal model for radiology report generation
Hyland, Stephanie L.
Bannur, Shruthi
Bouzid, Kenza
Castro, Daniel C.
Ranjit, Mercy
Schwaighofer, Anton
Pérez-García, Fernando
Salvatelli, Valentina
Srivastav, Shaury
Thieme, Anja
Codella, Noel
Lungren, Matthew P.
Wetscherek, Maria Teodora
Oktay, Ozan
Alvarez-Valle, Javier
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s) can be equipped with multimodal capabilities through alignment with pre-trained vision encoders. On natural images, this has been shown to allow multimodal models to gain image understanding and description capabilities. Our proposed model (MAIRA-1) leverages a CXR-specific image encoder in conjunction with a fine-tuned large language model based on Vicuna-7B, and text-based data augmentation, to produce reports with state-of-the-art quality. In particular, MAIRA-1 significantly improves on the radiologist-aligned RadCliQ metric and across all lexical metrics considered. Manual review of model outputs demonstrates promising fluency and accuracy of generated reports while uncovering failure modes not captured by existing evaluation practices. More information and resources can be found on the project website: https://aka.ms/maira.
title MAIRA-1: A specialised large multimodal model for radiology report generation
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.13668