MAIRA-1: A specialised large multimodal model for radiology report generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909183075418112 |
|---|---|
| author | Hyland, Stephanie L. Bannur, Shruthi Bouzid, Kenza Castro, Daniel C. Ranjit, Mercy Schwaighofer, Anton Pérez-García, Fernando Salvatelli, Valentina Srivastav, Shaury Thieme, Anja Codella, Noel Lungren, Matthew P. Wetscherek, Maria Teodora Oktay, Ozan Alvarez-Valle, Javier |
| author_facet | Hyland, Stephanie L. Bannur, Shruthi Bouzid, Kenza Castro, Daniel C. Ranjit, Mercy Schwaighofer, Anton Pérez-García, Fernando Salvatelli, Valentina Srivastav, Shaury Thieme, Anja Codella, Noel Lungren, Matthew P. Wetscherek, Maria Teodora Oktay, Ozan Alvarez-Valle, Javier |
| contents | We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s) can be equipped with multimodal capabilities through alignment with pre-trained vision encoders. On natural images, this has been shown to allow multimodal models to gain image understanding and description capabilities. Our proposed model (MAIRA-1) leverages a CXR-specific image encoder in conjunction with a fine-tuned large language model based on Vicuna-7B, and text-based data augmentation, to produce reports with state-of-the-art quality. In particular, MAIRA-1 significantly improves on the radiologist-aligned RadCliQ metric and across all lexical metrics considered. Manual review of model outputs demonstrates promising fluency and accuracy of generated reports while uncovering failure modes not captured by existing evaluation practices. More information and resources can be found on the project website: https://aka.ms/maira. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_13668 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | MAIRA-1: A specialised large multimodal model for radiology report generation Hyland, Stephanie L. Bannur, Shruthi Bouzid, Kenza Castro, Daniel C. Ranjit, Mercy Schwaighofer, Anton Pérez-García, Fernando Salvatelli, Valentina Srivastav, Shaury Thieme, Anja Codella, Noel Lungren, Matthew P. Wetscherek, Maria Teodora Oktay, Ozan Alvarez-Valle, Javier Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s) can be equipped with multimodal capabilities through alignment with pre-trained vision encoders. On natural images, this has been shown to allow multimodal models to gain image understanding and description capabilities. Our proposed model (MAIRA-1) leverages a CXR-specific image encoder in conjunction with a fine-tuned large language model based on Vicuna-7B, and text-based data augmentation, to produce reports with state-of-the-art quality. In particular, MAIRA-1 significantly improves on the radiologist-aligned RadCliQ metric and across all lexical metrics considered. Manual review of model outputs demonstrates promising fluency and accuracy of generated reports while uncovering failure modes not captured by existing evaluation practices. More information and resources can be found on the project website: https://aka.ms/maira. |
| title | MAIRA-1: A specialised large multimodal model for radiology report generation |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2311.13668 |