MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908940469534720 |
|---|---|
| author | Shi, Baorong Cui, Bo Jiang, Boyuan Yu, Deli Qian, Fang Yang, Haihua Wang, Huichao Chen, Jiale Pan, Jianfei Cao, Jieqiong Lin, Jinghao Wu, Kai Yang, Lin Yao, Shengsheng Chen, Tao Xiao, Xiaojun Ji, Xiaozhong Wang, Xu He, Yijun Yang, Zhixiong |
| author_facet | Shi, Baorong Cui, Bo Jiang, Boyuan Yu, Deli Qian, Fang Yang, Haihua Wang, Huichao Chen, Jiale Pan, Jianfei Cao, Jieqiong Lin, Jinghao Wu, Kai Yang, Lin Yao, Shengsheng Chen, Tao Xiao, Xiaojun Ji, Xiaozhong Wang, Xu He, Yijun Yang, Zhixiong |
| contents | We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world clinical applications. MedXIAOHE achieves state-of-the-art performance across diverse medical benchmarks and surpasses leading closed-source multimodal systems on multiple capabilities. To achieve this, we propose an entity-aware continual pretraining framework that organizes heterogeneous medical corpora to broaden knowledge coverage and reduce long-tail gaps (e.g., rare diseases). For medical expert-level reasoning and interaction, MedXIAOHE incorporates diverse medical reasoning patterns via reinforcement learning and tool-augmented agentic training, enabling multi-step diagnostic reasoning with verifiable decision traces. To improve reliability in real-world use, MedXIAOHE integrates user-preference rubrics, evidence-grounded reasoning, and low-hallucination long-form report generation, with improved adherence to medical instructions. We release this report to document our practical design choices, scaling insights, and evaluation framework, hoping to inspire further research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_12705 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Shi, Baorong Cui, Bo Jiang, Boyuan Yu, Deli Qian, Fang Yang, Haihua Wang, Huichao Chen, Jiale Pan, Jianfei Cao, Jieqiong Lin, Jinghao Wu, Kai Yang, Lin Yao, Shengsheng Chen, Tao Xiao, Xiaojun Ji, Xiaozhong Wang, Xu He, Yijun Yang, Zhixiong Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Image and Video Processing We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world clinical applications. MedXIAOHE achieves state-of-the-art performance across diverse medical benchmarks and surpasses leading closed-source multimodal systems on multiple capabilities. To achieve this, we propose an entity-aware continual pretraining framework that organizes heterogeneous medical corpora to broaden knowledge coverage and reduce long-tail gaps (e.g., rare diseases). For medical expert-level reasoning and interaction, MedXIAOHE incorporates diverse medical reasoning patterns via reinforcement learning and tool-augmented agentic training, enabling multi-step diagnostic reasoning with verifiable decision traces. To improve reliability in real-world use, MedXIAOHE integrates user-preference rubrics, evidence-grounded reasoning, and low-hallucination long-form report generation, with improved adherence to medical instructions. We release this report to document our practical design choices, scaling insights, and evaluation framework, hoping to inspire further research. |
| title | MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Image and Video Processing |
| url | https://arxiv.org/abs/2602.12705 |