LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Zhifan, Yang, Dong, Nath, Vishwesh, Parida, Abhijeet, Kulkarni, Nishad P., Xu, Ziyue, Xu, Daguang, Anwar, Syed Muhammad, Roth, Holger R., Linguraru, Marius George
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918354160189440
author Jiang, Zhifan
Yang, Dong
Nath, Vishwesh
Parida, Abhijeet
Kulkarni, Nishad P.
Xu, Ziyue
Xu, Daguang
Anwar, Syed Muhammad
Roth, Holger R.
Linguraru, Marius George
author_facet Jiang, Zhifan
Yang, Dong
Nath, Vishwesh
Parida, Abhijeet
Kulkarni, Nishad P.
Xu, Ziyue
Xu, Daguang
Anwar, Syed Muhammad
Roth, Holger R.
Linguraru, Marius George
contents Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting radiologists in decision-making by the analysis of radiology imaging data such as chest X-rays (CXR) via a visual and natural language question-answering (VQA) interface. When longitudinal imaging is available, radiologists analyze temporal changes, which are essential for accurate diagnosis and prognosis. The manual longitudinal analysis is a time-consuming process, motivating the development of a training framework that can provide prognostic capabilities. We introduce a novel training framework LUMEN, that is optimized for longitudinal CXR interpretation, leveraging multi-image and multi-task instruction fine-tuning to enhance prognostic and diagnostic performance. We conduct experiments on the publicly available MIMIC-CXR and its associated Medical-Diff-VQA datasets. We further formulate and construct a novel instruction-following dataset incorporating longitudinal studies, enabling the development of a prognostic VQA task. Our method demonstrates significant improvements over baseline models in diagnostic VQA tasks, and more importantly, shows promising potential for prognostic capabilities. These results underscore the value of well-designed, instruction-tuned VLMs in enabling more accurate and clinically meaningful radiological interpretation of longitudinal radiological imaging data.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21142
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis
Jiang, Zhifan
Yang, Dong
Nath, Vishwesh
Parida, Abhijeet
Kulkarni, Nishad P.
Xu, Ziyue
Xu, Daguang
Anwar, Syed Muhammad
Roth, Holger R.
Linguraru, Marius George
Computer Vision and Pattern Recognition
Machine Learning
Large vision-language models (VLMs) have evolved from general-purpose applications to specialized use cases such as in the clinical domain, demonstrating potential for decision support in radiology. One promising application is assisting radiologists in decision-making by the analysis of radiology imaging data such as chest X-rays (CXR) via a visual and natural language question-answering (VQA) interface. When longitudinal imaging is available, radiologists analyze temporal changes, which are essential for accurate diagnosis and prognosis. The manual longitudinal analysis is a time-consuming process, motivating the development of a training framework that can provide prognostic capabilities. We introduce a novel training framework LUMEN, that is optimized for longitudinal CXR interpretation, leveraging multi-image and multi-task instruction fine-tuning to enhance prognostic and diagnostic performance. We conduct experiments on the publicly available MIMIC-CXR and its associated Medical-Diff-VQA datasets. We further formulate and construct a novel instruction-following dataset incorporating longitudinal studies, enabling the development of a prognostic VQA task. Our method demonstrates significant improvements over baseline models in diagnostic VQA tasks, and more importantly, shows promising potential for prognostic capabilities. These results underscore the value of well-designed, instruction-tuned VLMs in enabling more accurate and clinically meaningful radiological interpretation of longitudinal radiological imaging data.
title LUMEN: Longitudinal Multi-Modal Radiology Model for Prognosis and Diagnosis
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2602.21142