MediSee: Reasoning-based Pixel-level Perception in Medical Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tong, Qinyue, Lu, Ziqian, Liu, Jun, Zheng, Yangming, Lu, Zheming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917995978162176
author Tong, Qinyue
Lu, Ziqian
Liu, Jun
Zheng, Yangming
Lu, Zheming
author_facet Tong, Qinyue
Lu, Ziqian
Liu, Jun
Zheng, Yangming
Lu, Zheming
contents Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge required for input is a huge obstacle for general public, which greatly reduces the universality of these methods. Compared with these domain-specialized auxiliary information, general users tend to rely on oral queries that require logical reasoning. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for medical reasoning segmentation and detection. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11008
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MediSee: Reasoning-based Pixel-level Perception in Medical Images
Tong, Qinyue
Lu, Ziqian
Liu, Jun
Zheng, Yangming
Lu, Zheming
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite remarkable advancements in pixel-level medical image perception, existing methods are either limited to specific tasks or heavily rely on accurate bounding boxes or text labels as input prompts. However, the medical knowledge required for input is a huge obstacle for general public, which greatly reduces the universality of these methods. Compared with these domain-specialized auxiliary information, general users tend to rely on oral queries that require logical reasoning. In this paper, we introduce a novel medical vision task: Medical Reasoning Segmentation and Detection (MedSD), which aims to comprehend implicit queries about medical images and generate the corresponding segmentation mask and bounding box for the target object. To accomplish this task, we first introduce a Multi-perspective, Logic-driven Medical Reasoning Segmentation and Detection (MLMR-SD) dataset, which encompasses a substantial collection of medical entity targets along with their corresponding reasoning. Furthermore, we propose MediSee, an effective baseline model designed for medical reasoning segmentation and detection. The experimental results indicate that the proposed method can effectively address MedSD with implicit colloquial queries and outperform traditional medical referring segmentation methods.
title MediSee: Reasoning-based Pixel-level Perception in Medical Images
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2504.11008