MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Baorong, Cui, Bo, Jiang, Boyuan, Yu, Deli, Qian, Fang, Yang, Haihua, Wang, Huichao, Chen, Jiale, Pan, Jianfei, Cao, Jieqiong, Lin, Jinghao, Wu, Kai, Yang, Lin, Yao, Shengsheng, Chen, Tao, Xiao, Xiaojun, Ji, Xiaozhong, Wang, Xu, He, Yijun, Yang, Zhixiong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs
by: Lin, Jianghang, et al.
Published: (2026)
by: Lin, Jianghang, et al.
Published: (2026)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026)
by: Yang, Lin, et al.
Published: (2026)
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
MedEyes: Learning Dynamic Visual Focus for Medical Progressive Diagnosis
by: Zhu, Chunzheng, et al.
Published: (2025)
by: Zhu, Chunzheng, et al.
Published: (2025)
ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following
by: Yang, Yuancheng, et al.
Published: (2026)
by: Yang, Yuancheng, et al.
Published: (2026)
CloudFort: Enhancing Robustness of 3D Point Cloud Classification Against Backdoor Attacks via Spatial Partitioning and Ensemble Prediction
by: Lan, Wenhao, et al.
Published: (2024)
by: Lan, Wenhao, et al.
Published: (2024)
InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
Growing Visual Generative Capacity for Pre-Trained MLLMs
by: Wang, Hanyu, et al.
Published: (2025)
by: Wang, Hanyu, et al.
Published: (2025)
MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
by: Liu, Jiyao, et al.
Published: (2025)
by: Liu, Jiyao, et al.
Published: (2025)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
by: Bao, Zhijie, et al.
Published: (2026)
by: Bao, Zhijie, et al.
Published: (2026)
VLANeXt: Recipes for Building Strong VLA Models
by: Wu, Xiao-Ming, et al.
Published: (2026)
by: Wu, Xiao-Ming, et al.
Published: (2026)
AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts
by: Huang, Baorong, et al.
Published: (2025)
by: Huang, Baorong, et al.
Published: (2025)
LATA: A Tool for LLM-Assisted Translation Annotation
by: Huang, Baorong, et al.
Published: (2026)
by: Huang, Baorong, et al.
Published: (2026)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
X-Fi: A Modality-Invariant Foundation Model for Multimodal Human Sensing
by: Chen, Xinyan, et al.
Published: (2024)
by: Chen, Xinyan, et al.
Published: (2024)
AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
by: Chang, Boyu, et al.
Published: (2026)
by: Chang, Boyu, et al.
Published: (2026)
OpenView: Empowering MLLMs with Out-of-view VQA
by: Chen, Qixiang, et al.
Published: (2025)
by: Chen, Qixiang, et al.
Published: (2025)
The Forgotten Shield: Safety Grafting in Parameter-Space for Medical MLLMs
by: Zhao, Jiale, et al.
Published: (2025)
by: Zhao, Jiale, et al.
Published: (2025)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
RMAvatar: Photorealistic Human Avatar Reconstruction from Monocular Video Based on Rectified Mesh-embedded Gaussians
by: Peng, Sen, et al.
Published: (2025)
by: Peng, Sen, et al.
Published: (2025)
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
by: Zhu, Chunzheng, et al.
Published: (2026)
by: Zhu, Chunzheng, et al.
Published: (2026)
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning
by: Chen, Guizhen, et al.
Published: (2025)
by: Chen, Guizhen, et al.
Published: (2025)
Towards Codable Watermarking for Injecting Multi-bits Information to LLMs
by: Wang, Lean, et al.
Published: (2023)
by: Wang, Lean, et al.
Published: (2023)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
by: Yu, Tianyu, et al.
Published: (2025)
by: Yu, Tianyu, et al.
Published: (2025)
GRIT: Teaching MLLMs to Think with Images
by: Fan, Yue, et al.
Published: (2025)
by: Fan, Yue, et al.
Published: (2025)
RynnEC: Bringing MLLMs into Embodied World
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
by: Wang, Serene, et al.
Published: (2026)
by: Wang, Serene, et al.
Published: (2026)
Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry
by: Lan, Wenhao, et al.
Published: (2026)
by: Lan, Wenhao, et al.
Published: (2026)
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning
by: Wu, Yang, et al.
Published: (2024)
by: Wu, Yang, et al.
Published: (2024)
Proposal Report for the 2nd SciCAP Competition 2024
by: Li, Pengpeng, et al.
Published: (2024)
by: Li, Pengpeng, et al.
Published: (2024)
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
by: Shang, Fangxin, et al.
Published: (2025)
by: Shang, Fangxin, et al.
Published: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
by: Du, Yuetian, et al.
Published: (2026)
by: Du, Yuetian, et al.
Published: (2026)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
by: Li, Sunzhu, et al.
Published: (2025)
by: Li, Sunzhu, et al.
Published: (2025)
MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
by: Du, Yexing, et al.
Published: (2025)
by: Du, Yexing, et al.
Published: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
by: Chen, Kun, et al.
Published: (2025)
by: Chen, Kun, et al.
Published: (2025)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
by: Liu, Yilian, et al.
Published: (2026)
by: Liu, Yilian, et al.
Published: (2026)
Similar Items
-
Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs
by: Lin, Jianghang, et al.
Published: (2026) -
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
by: Yang, Lin, et al.
Published: (2026) -
DEEPMED: Building a Medical DeepResearch Agent via Multi-hop Med-Search Data and Turn-Controlled Agentic Training & Inference
by: Wang, Zihan, et al.
Published: (2026) -
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025) -
MedEyes: Learning Dynamic Visual Focus for Medical Progressive Diagnosis
by: Zhu, Chunzheng, et al.
Published: (2025)