MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yu, Peng, Zelin, Zhao, Yichen, Yang, Piao, Yang, Xiaokang, Shen, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916791876321280
author Huang, Yu
Peng, Zelin
Zhao, Yichen
Yang, Piao
Yang, Xiaokang
Shen, Wei
author_facet Huang, Yu
Peng, Zelin
Zhao, Yichen
Yang, Piao
Yang, Xiaokang
Shen, Wei
contents Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabilities to understand complex clinical questions. While recent advancements in multimodal large language models (MLLMs) have improved medical question-answering (QA) tasks, most methods struggle to generate precise segmentation masks, limiting their application in automatic medical diagnosis. In this paper, we introduce medical image reasoning segmentation, a novel task that aims to generate segmentation masks based on complex and implicit medical instructions. To address this, we propose MedSeg-R, an end-to-end framework that leverages the reasoning abilities of MLLMs to interpret clinical questions while also capable of producing corresponding precise segmentation masks for medical images. It is built on two core components: 1) a global context understanding module that interprets images and comprehends complex medical instructions to generate multi-modal intermediate tokens, and 2) a pixel-level grounding module that decodes these tokens to produce precise segmentation masks and textual responses. Furthermore, we introduce MedSeg-QA, a large-scale dataset tailored for the medical image reasoning segmentation task. It includes over 10,000 image-mask pairs and multi-turn conversations, automatically annotated using large language models and refined through physician reviews. Experiments show MedSeg-R's superior performance across several benchmarks, achieving high segmentation accuracy and enabling interpretable textual analysis of medical images.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10465
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Huang, Yu
Peng, Zelin
Zhao, Yichen
Yang, Piao
Yang, Xiaokang
Shen, Wei
Computer Vision and Pattern Recognition
Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabilities to understand complex clinical questions. While recent advancements in multimodal large language models (MLLMs) have improved medical question-answering (QA) tasks, most methods struggle to generate precise segmentation masks, limiting their application in automatic medical diagnosis. In this paper, we introduce medical image reasoning segmentation, a novel task that aims to generate segmentation masks based on complex and implicit medical instructions. To address this, we propose MedSeg-R, an end-to-end framework that leverages the reasoning abilities of MLLMs to interpret clinical questions while also capable of producing corresponding precise segmentation masks for medical images. It is built on two core components: 1) a global context understanding module that interprets images and comprehends complex medical instructions to generate multi-modal intermediate tokens, and 2) a pixel-level grounding module that decodes these tokens to produce precise segmentation masks and textual responses. Furthermore, we introduce MedSeg-QA, a large-scale dataset tailored for the medical image reasoning segmentation task. It includes over 10,000 image-mask pairs and multi-turn conversations, automatically annotated using large language models and refined through physician reviews. Experiments show MedSeg-R's superior performance across several benchmarks, achieving high segmentation accuracy and enabling interpretable textual analysis of medical images.
title MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.10465