Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Zhongzhen, Mu, Linjie, Zhu, Yakun, Zhao, Xiangyu, Zhang, Shaoting, Zhang, Xiaofan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909881368313856
author Huang, Zhongzhen
Mu, Linjie
Zhu, Yakun
Zhao, Xiangyu
Zhang, Shaoting
Zhang, Xiaofan
author_facet Huang, Zhongzhen
Mu, Linjie
Zhu, Yakun
Zhao, Xiangyu
Zhang, Shaoting
Zhang, Xiaofan
contents Effective clinical decision-making depends on iterative, multimodal reasoning across diverse sources of evidence. The recent emergence of multimodal reasoning models has significantly transformed the landscape of solving complex tasks. Although such models have achieved notable success in mathematics and science, their application to medical domains remains underexplored. In this work, we propose \textit{MedE$^2$}, a two-stage post-training pipeline that elicits and then enhances multimodal reasoning for medical domains. In Stage-I, we fine-tune models using 2,000 text-only data samples containing precisely orchestrated reasoning demonstrations to elicit reasoning behaviors. In Stage-II, we further enhance the model's reasoning capabilities using 1,500 rigorously curated multimodal medical cases, aligning model reasoning outputs with our proposed multimodal medical reasoning preference. Extensive experiments demonstrate the efficacy and reliability of \textit{MedE$^2$} in improving the reasoning performance of medical multimodal models. Notably, models trained with \textit{MedE$^2$} consistently outperform baselines across multiple medical multimodal benchmarks. Additional validation on larger models and under inference-time scaling further confirms the robustness and practical utility of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
Huang, Zhongzhen
Mu, Linjie
Zhu, Yakun
Zhao, Xiangyu
Zhang, Shaoting
Zhang, Xiaofan
Computation and Language
Artificial Intelligence
Effective clinical decision-making depends on iterative, multimodal reasoning across diverse sources of evidence. The recent emergence of multimodal reasoning models has significantly transformed the landscape of solving complex tasks. Although such models have achieved notable success in mathematics and science, their application to medical domains remains underexplored. In this work, we propose \textit{MedE$^2$}, a two-stage post-training pipeline that elicits and then enhances multimodal reasoning for medical domains. In Stage-I, we fine-tune models using 2,000 text-only data samples containing precisely orchestrated reasoning demonstrations to elicit reasoning behaviors. In Stage-II, we further enhance the model's reasoning capabilities using 1,500 rigorously curated multimodal medical cases, aligning model reasoning outputs with our proposed multimodal medical reasoning preference. Extensive experiments demonstrate the efficacy and reliability of \textit{MedE$^2$} in improving the reasoning performance of medical multimodal models. Notably, models trained with \textit{MedE$^2$} consistently outperform baselines across multiple medical multimodal benchmarks. Additional validation on larger models and under inference-time scaling further confirms the robustness and practical utility of our approach.
title Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.23118