O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Huang, Zhongzhen, Geng, Gui, Hua, Shengyi, Huang, Zhen, Zou, Haoyang, Zhang, Shaoting, Liu, Pengfei, Zhang, Xiaofan
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913644214747136
author Huang, Zhongzhen
Geng, Gui
Hua, Shengyi
Huang, Zhen
Zou, Haoyang
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
author_facet Huang, Zhongzhen
Geng, Gui
Hua, Shengyi
Huang, Zhen
Zou, Haoyang
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
contents Building upon our previous investigations of O1 replication (Part 1: Journey Learning [Qin et al., 2024] and Part 2: Distillation [Huang et al., 2024]), this work explores the potential of inference-time scaling in large language models (LLMs) for medical reasoning tasks, ranging from diagnostic decision-making to treatment planning. Through extensive experiments on medical benchmarks of varying complexity (MedQA, Medbullets, and JAMA Clinical Challenges), our investigation reveals several key insights: (1) Increasing inference time does lead to improved performance. With a modest training set of 500 samples, our model yields substantial performance improvements of 6%-11%. (2) Task complexity directly correlates with the required length of reasoning chains, confirming the necessity of extended thought processes for challenging problems. (3) The differential diagnoses generated by our model adhere to the principles of the hypothetico-deductive method, producing a list of potential conditions that may explain a patient's symptoms and systematically narrowing these possibilities by evaluating the evidence. These findings demonstrate the promising synergy between inference-time scaling and journey learning in advancing LLMs' real-world clinical reasoning capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2501_06458
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
Huang, Zhongzhen
Geng, Gui
Hua, Shengyi
Huang, Zhen
Zou, Haoyang
Zhang, Shaoting
Liu, Pengfei
Zhang, Xiaofan
Computation and Language
Building upon our previous investigations of O1 replication (Part 1: Journey Learning [Qin et al., 2024] and Part 2: Distillation [Huang et al., 2024]), this work explores the potential of inference-time scaling in large language models (LLMs) for medical reasoning tasks, ranging from diagnostic decision-making to treatment planning. Through extensive experiments on medical benchmarks of varying complexity (MedQA, Medbullets, and JAMA Clinical Challenges), our investigation reveals several key insights: (1) Increasing inference time does lead to improved performance. With a modest training set of 500 samples, our model yields substantial performance improvements of 6%-11%. (2) Task complexity directly correlates with the required length of reasoning chains, confirming the necessity of extended thought processes for challenging problems. (3) The differential diagnoses generated by our model adhere to the principles of the hypothetico-deductive method, producing a list of potential conditions that may explain a patient's symptoms and systematically narrowing these possibilities by evaluating the evidence. These findings demonstrate the promising synergy between inference-time scaling and journey learning in advancing LLMs' real-world clinical reasoning capabilities.
title O1 Replication Journey -- Part 3: Inference-time Scaling for Medical Reasoning
topic Computation and Language
url https://arxiv.org/abs/2501.06458