Text Embedded Swin-UMamba for DeepLesion Segmentation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cheng, Ruida, Mathai, Tejas Sudharshan, Mukherjee, Pritam, Hou, Benjamin, Zhu, Qingqing, Lu, Zhiyong, McAuliffe, Matthew, Summers, Ronald M.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908713935175680
author Cheng, Ruida
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Hou, Benjamin
Zhu, Qingqing
Lu, Zhiyong
McAuliffe, Matthew
Summers, Ronald M.
author_facet Cheng, Ruida
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Hou, Benjamin
Zhu, Qingqing
Lu, Zhiyong
McAuliffe, Matthew
Summers, Ronald M.
contents Segmentation of lesions on CT enables automatic measurement for clinical assessment of chronic diseases (e.g., lymphoma). Integrating large language models (LLMs) into the lesion segmentation workflow has the potential to combine imaging features with descriptions of lesion characteristics from the radiology reports. In this study, we investigate the feasibility of integrating text into the Swin-UMamba architecture for the task of lesion segmentation. The publicly available ULS23 DeepLesion dataset was used along with short-form descriptions of the findings from the reports. On the test dataset, our method achieved a high Dice score of 82.64, and a low Hausdorff distance of 6.34 pixels was obtained for lesion segmentation. The proposed Text-Swin-U/Mamba model outperformed prior approaches: 37.79% improvement over the LLM-driven LanGuideMedSeg model (p < 0.001), and surpassed the purely image-based XLSTM-UNet and nnUNet models by 2.58% and 1.01%, respectively. The dataset and code can be accessed at https://github.com/ruida/LLM-Swin-UMamba
format Preprint
id arxiv_https___arxiv_org_abs_2508_06453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text Embedded Swin-UMamba for DeepLesion Segmentation
Cheng, Ruida
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Hou, Benjamin
Zhu, Qingqing
Lu, Zhiyong
McAuliffe, Matthew
Summers, Ronald M.
Computer Vision and Pattern Recognition
Artificial Intelligence
Segmentation of lesions on CT enables automatic measurement for clinical assessment of chronic diseases (e.g., lymphoma). Integrating large language models (LLMs) into the lesion segmentation workflow has the potential to combine imaging features with descriptions of lesion characteristics from the radiology reports. In this study, we investigate the feasibility of integrating text into the Swin-UMamba architecture for the task of lesion segmentation. The publicly available ULS23 DeepLesion dataset was used along with short-form descriptions of the findings from the reports. On the test dataset, our method achieved a high Dice score of 82.64, and a low Hausdorff distance of 6.34 pixels was obtained for lesion segmentation. The proposed Text-Swin-U/Mamba model outperformed prior approaches: 37.79% improvement over the LLM-driven LanGuideMedSeg model (p < 0.001), and surpassed the purely image-based XLSTM-UNet and nnUNet models by 2.58% and 1.01%, respectively. The dataset and code can be accessed at https://github.com/ruida/LLM-Swin-UMamba
title Text Embedded Swin-UMamba for DeepLesion Segmentation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.06453