R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Xiao, Li, Yuehang, Wang, Fuling, Wang, Shiao, Li, Chuanfu, Jiang, Bo
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910034334580736
author Wang, Xiao
Li, Yuehang
Wang, Fuling
Wang, Shiao
Li, Chuanfu
Jiang, Bo
author_facet Wang, Xiao
Li, Yuehang
Wang, Fuling
Wang, Shiao
Li, Chuanfu
Jiang, Bo
contents Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. They usually adopt a Transformer to extract the visual features of a given X-ray image, and then, feed them into the LLM for text generation. How to extract more effective information for the LLMs to help them improve final results is an urgent problem that needs to be solved. Additionally, the use of visual Transformer models also brings high computational complexity. To address these issues, this paper proposes a novel context-guided efficient radiology report generation framework. Specifically, we introduce the Mamba as the vision backbone with linear complexity, and the performance obtained is comparable to that of the strong Transformer model. More importantly, we perform context retrieval from the training set for samples within each mini-batch during the training phase, utilizing both positively and negatively related samples to enhance feature representation and discriminative learning. Subsequently, we feed the vision tokens, context information, and prompt statements to invoke the LLM for generating high-quality medical reports. Extensive experiments on three X-ray report generation datasets (i.e., IU X-Ray, MIMIC-CXR, CheXpert Plus) fully validated the effectiveness of our proposed model. The source code is available at https://github.com/Event-AHU/Medical_Image_Analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09743
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation
Wang, Xiao
Li, Yuehang
Wang, Fuling
Wang, Shiao
Li, Chuanfu
Jiang, Bo
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Inspired by the tremendous success of Large Language Models (LLMs), existing Radiology report generation methods attempt to leverage large models to achieve better performance. They usually adopt a Transformer to extract the visual features of a given X-ray image, and then, feed them into the LLM for text generation. How to extract more effective information for the LLMs to help them improve final results is an urgent problem that needs to be solved. Additionally, the use of visual Transformer models also brings high computational complexity. To address these issues, this paper proposes a novel context-guided efficient radiology report generation framework. Specifically, we introduce the Mamba as the vision backbone with linear complexity, and the performance obtained is comparable to that of the strong Transformer model. More importantly, we perform context retrieval from the training set for samples within each mini-batch during the training phase, utilizing both positively and negatively related samples to enhance feature representation and discriminative learning. Subsequently, we feed the vision tokens, context information, and prompt statements to invoke the LLM for generating high-quality medical reports. Extensive experiments on three X-ray report generation datasets (i.e., IU X-Ray, MIMIC-CXR, CheXpert Plus) fully validated the effectiveness of our proposed model. The source code is available at https://github.com/Event-AHU/Medical_Image_Analysis.
title R2GenCSR: Mining Contextual and Residual Information for LLMs-based Radiology Report Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2408.09743