RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Chengrun, Royer, Corentin, Luo, Haozhe, Wittmann, Bastian, Li, Xia, Hamamci, Ibrahim, Er, Sezgin, Sekuboyina, Anjany, Menze, Bjoern
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909860818321408
author Li, Chengrun
Royer, Corentin
Luo, Haozhe
Wittmann, Bastian
Li, Xia
Hamamci, Ibrahim
Er, Sezgin
Sekuboyina, Anjany
Menze, Bjoern
author_facet Li, Chengrun
Royer, Corentin
Luo, Haozhe
Wittmann, Bastian
Li, Xia
Hamamci, Ibrahim
Er, Sezgin
Sekuboyina, Anjany
Menze, Bjoern
contents Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as assistive systems that fail to provide both modalities simultaneously offer limited value to medical practitioners. To alleviate this limitation, we first introduce RadDiagSeg-D, a dataset combining abnormality detection, diagnosis, and multi-target segmentation into a unified and hierarchical task. RadDiagSeg-D covers multiple imaging modalities and is precisely designed to support the development of models that produce descriptive text and corresponding segmentation masks in tandem. Subsequently, we leverage the dataset to propose a novel vision-language model, RadDiagSeg-M, capable of joint abnormality detection, diagnosis, and flexible segmentation. RadDiagSeg-M provides highly informative and clinically useful outputs, effectively addressing the need to enrich contextual information for assistive diagnosis. Finally, we benchmark RadDiagSeg-M and showcase its strong performance across all components involved in the task of multi-target text-and-mask generation, establishing a robust and competitive baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
Li, Chengrun
Royer, Corentin
Luo, Haozhe
Wittmann, Bastian
Li, Xia
Hamamci, Ibrahim
Er, Sezgin
Sekuboyina, Anjany
Menze, Bjoern
Computer Vision and Pattern Recognition
Artificial Intelligence
68
I.4.6
Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as assistive systems that fail to provide both modalities simultaneously offer limited value to medical practitioners. To alleviate this limitation, we first introduce RadDiagSeg-D, a dataset combining abnormality detection, diagnosis, and multi-target segmentation into a unified and hierarchical task. RadDiagSeg-D covers multiple imaging modalities and is precisely designed to support the development of models that produce descriptive text and corresponding segmentation masks in tandem. Subsequently, we leverage the dataset to propose a novel vision-language model, RadDiagSeg-M, capable of joint abnormality detection, diagnosis, and flexible segmentation. RadDiagSeg-M provides highly informative and clinically useful outputs, effectively addressing the need to enrich contextual information for assistive diagnosis. Finally, we benchmark RadDiagSeg-M and showcase its strong performance across all components involved in the task of multi-target text-and-mask generation, establishing a robust and competitive baseline.
title RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
topic Computer Vision and Pattern Recognition
Artificial Intelligence
68
I.4.6
url https://arxiv.org/abs/2510.18188