MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yan, Zhonghao, Diao, Muxi, Yang, Yuxuan, Jing, Ruoyan, Xu, Jiayuan, Zhang, Kaizhou, Yang, Lele, Liu, Yanxi, Liang, Kongming, Ma, Zhanyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918343438499840
author Yan, Zhonghao
Diao, Muxi
Yang, Yuxuan
Jing, Ruoyan
Xu, Jiayuan
Zhang, Kaizhou
Yang, Lele
Liu, Yanxi
Liang, Kongming
Ma, Zhanyu
author_facet Yan, Zhonghao
Diao, Muxi
Yang, Yuxuan
Jing, Ruoyan
Xu, Jiayuan
Zhang, Kaizhou
Yang, Lele
Liu, Yanxi
Liang, Kongming
Ma, Zhanyu
contents Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision-language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08177
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
Yan, Zhonghao
Diao, Muxi
Yang, Yuxuan
Jing, Ruoyan
Xu, Jiayuan
Zhang, Kaizhou
Yang, Lele
Liu, Yanxi
Liang, Kongming
Ma, Zhanyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Accurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision-language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding.
title MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.08177