SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Jinge, Kim, Yunsoo, Shi, Daqian, Cliffton, David, Liu, Fenglin, Wu, Honghan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913509753749504
author Wu, Jinge
Kim, Yunsoo
Shi, Daqian
Cliffton, David
Liu, Fenglin
Wu, Honghan
author_facet Wu, Jinge
Kim, Yunsoo
Shi, Daqian
Cliffton, David
Liu, Fenglin
Wu, Honghan
contents Inspired by the success of large language models (LLMs), there is growing research interest in developing LLMs in the medical domain to assist clinicians. However, for hospitals, using closed-source commercial LLMs involves privacy issues, and developing open-source public LLMs requires large-scale computational resources, which are usually limited, especially in resource-efficient regions and low-income countries. We propose an open-source Small Language and Vision Assistant (SLaVA-CXR) that can be used for Chest X-Ray report automation. To efficiently train a small assistant, we first propose the Re$^3$Training method, which simulates the cognitive development of radiologists and optimizes the model in the Recognition, Reasoning, and Reporting training manner. Then, we introduce a data synthesis method, RADEX, which can generate a high-quality and diverse training corpus with privacy regulation compliance. The extensive experiments show that our SLaVA-CXR built on a 2.7B backbone not only outperforms but also achieves 6 times faster inference efficiency than previous state-of-the-art larger models.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13321
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation
Wu, Jinge
Kim, Yunsoo
Shi, Daqian
Cliffton, David
Liu, Fenglin
Wu, Honghan
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Inspired by the success of large language models (LLMs), there is growing research interest in developing LLMs in the medical domain to assist clinicians. However, for hospitals, using closed-source commercial LLMs involves privacy issues, and developing open-source public LLMs requires large-scale computational resources, which are usually limited, especially in resource-efficient regions and low-income countries. We propose an open-source Small Language and Vision Assistant (SLaVA-CXR) that can be used for Chest X-Ray report automation. To efficiently train a small assistant, we first propose the Re$^3$Training method, which simulates the cognitive development of radiologists and optimizes the model in the Recognition, Reasoning, and Reporting training manner. Then, we introduce a data synthesis method, RADEX, which can generate a high-quality and diverse training corpus with privacy regulation compliance. The extensive experiments show that our SLaVA-CXR built on a 2.7B backbone not only outperforms but also achieves 6 times faster inference efficiency than previous state-of-the-art larger models.
title SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.13321