OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ossowski, Timothy, Zhang, Sheng, Liu, Qianchu, Qin, Guanghui, Tan, Reuben, Naumann, Tristan, Hu, Junjie, Poon, Hoifung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915643857567744
author Ossowski, Timothy
Zhang, Sheng
Liu, Qianchu
Qin, Guanghui
Tan, Reuben
Naumann, Tristan
Hu, Junjie
Poon, Hoifung
author_facet Ossowski, Timothy
Zhang, Sheng
Liu, Qianchu
Qin, Guanghui
Tan, Reuben
Naumann, Tristan
Hu, Junjie
Poon, Hoifung
contents High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medical domain. Our work focuses on supervised fine-tuning (SFT) and explores data recipes that leverage structured reasoning traces. Using our proposed data recipe, we scale experiments to a dataset of over 8 million examples and 6.8 billion response tokens, achieving state-of-the-art performance among open-source models across diverse out-of-distribution medical benchmark tasks. Our results further indicate that curating a high-quality, diverse training dataset with varying structured reasoning trace lengths enables the fine-tuned model to self-calibrate its reasoning trajectory lengths based on the downstream task, without explicit supervision. We present key insights, describe the data curation strategy, and outline next steps toward developing robust medical vision-language reasoning system.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23269
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
Ossowski, Timothy
Zhang, Sheng
Liu, Qianchu
Qin, Guanghui
Tan, Reuben
Naumann, Tristan
Hu, Junjie
Poon, Hoifung
Artificial Intelligence
High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medical domain. Our work focuses on supervised fine-tuning (SFT) and explores data recipes that leverage structured reasoning traces. Using our proposed data recipe, we scale experiments to a dataset of over 8 million examples and 6.8 billion response tokens, achieving state-of-the-art performance among open-source models across diverse out-of-distribution medical benchmark tasks. Our results further indicate that curating a high-quality, diverse training dataset with varying structured reasoning trace lengths enables the fine-tuned model to self-calibrate its reasoning trajectory lengths based on the downstream task, without explicit supervision. We present key insights, describe the data curation strategy, and outline next steps toward developing robust medical vision-language reasoning system.
title OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
topic Artificial Intelligence
url https://arxiv.org/abs/2511.23269