Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Chunlei, Hou, Jingyang, Shi, Yilei, Hu, Jingliang, Zhu, Xiao Xiang, Mou, Lichao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916799412436992
author Li, Chunlei
Hou, Jingyang
Shi, Yilei
Hu, Jingliang
Zhu, Xiao Xiang
Mou, Lichao
author_facet Li, Chunlei
Hou, Jingyang
Shi, Yilei
Hu, Jingliang
Zhu, Xiao Xiang
Mou, Lichao
contents Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still deserves in-depth exploration. In this paper, we present MRG-LLM, a novel multimodal large language model (MLLM) that combines a frozen LLM with a learnable visual encoder and introduces a dynamic prompt customization mechanism. Our key innovation lies in generating instance-specific prompts tailored to individual medical images through conditional affine transformations derived from visual features. We propose two implementations: prompt-wise and promptbook-wise customization, enabling precise and targeted report generation. Extensive experiments on IU X-ray and MIMIC-CXR datasets demonstrate that MRG-LLM achieves state-of-the-art performance in medical report generation. Our code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2506_15477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
Li, Chunlei
Hou, Jingyang
Shi, Yilei
Hu, Jingliang
Zhu, Xiao Xiang
Mou, Lichao
Computer Vision and Pattern Recognition
Medical report generation from imaging data remains a challenging task in clinical practice. While large language models (LLMs) show great promise in addressing this challenge, their effective integration with medical imaging data still deserves in-depth exploration. In this paper, we present MRG-LLM, a novel multimodal large language model (MLLM) that combines a frozen LLM with a learnable visual encoder and introduces a dynamic prompt customization mechanism. Our key innovation lies in generating instance-specific prompts tailored to individual medical images through conditional affine transformations derived from visual features. We propose two implementations: prompt-wise and promptbook-wise customization, enabling precise and targeted report generation. Extensive experiments on IU X-ray and MIMIC-CXR datasets demonstrate that MRG-LLM achieves state-of-the-art performance in medical report generation. Our code will be made publicly available.
title Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.15477