Disentangling Reasoning and Knowledge in Medical Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Thapa, Rahul, Wu, Qingyang, Wu, Kevin, Zhang, Harrison, Zhang, Angela, Wu, Eric, Ye, Haotian, Bedi, Suhana, Aresh, Nevin, Boen, Joseph, Reddy, Shriya, Athiwaratkun, Ben, Song, Shuaiwen Leon, Zou, James
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918068821688320
author Thapa, Rahul
Wu, Qingyang
Wu, Kevin
Zhang, Harrison
Zhang, Angela
Wu, Eric
Ye, Haotian
Bedi, Suhana
Aresh, Nevin
Boen, Joseph
Reddy, Shriya
Athiwaratkun, Ben
Song, Shuaiwen Leon
Zou, James
author_facet Thapa, Rahul
Wu, Qingyang
Wu, Kevin
Zhang, Harrison
Zhang, Angela
Wu, Eric
Ye, Haotian
Bedi, Suhana
Aresh, Nevin
Boen, Joseph
Reddy, Shriya
Athiwaratkun, Ben
Song, Shuaiwen Leon
Zou, James
contents Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall. We address this by separating 11 biomedical QA benchmarks into reasoning- and knowledge-focused subsets using a PubMedBERT classifier that reaches 81 percent accuracy, comparable to human performance. Our analysis shows that only 32.8 percent of questions require complex reasoning. We evaluate biomedical models (HuatuoGPT-o1, MedReason, m1) and general-domain models (DeepSeek-R1, o4-mini, Qwen3), finding consistent gaps between knowledge and reasoning performance. For example, HuatuoGPT-o1 scores 56.9 on knowledge but only 44.8 on reasoning. In adversarial tests where models are misled with incorrect initial reasoning, biomedical models degrade sharply, while larger or RL-trained general models show more robustness. To address this, we train BioMed-R1 using fine-tuning and reinforcement learning on reasoning-heavy examples. It achieves the strongest performance among similarly sized models. Further gains may come from incorporating clinical case reports and training with adversarial and backtracking scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11462
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Disentangling Reasoning and Knowledge in Medical Large Language Models
Thapa, Rahul
Wu, Qingyang
Wu, Kevin
Zhang, Harrison
Zhang, Angela
Wu, Eric
Ye, Haotian
Bedi, Suhana
Aresh, Nevin
Boen, Joseph
Reddy, Shriya
Athiwaratkun, Ben
Song, Shuaiwen Leon
Zou, James
Computation and Language
Artificial Intelligence
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reasoning with factual recall. We address this by separating 11 biomedical QA benchmarks into reasoning- and knowledge-focused subsets using a PubMedBERT classifier that reaches 81 percent accuracy, comparable to human performance. Our analysis shows that only 32.8 percent of questions require complex reasoning. We evaluate biomedical models (HuatuoGPT-o1, MedReason, m1) and general-domain models (DeepSeek-R1, o4-mini, Qwen3), finding consistent gaps between knowledge and reasoning performance. For example, HuatuoGPT-o1 scores 56.9 on knowledge but only 44.8 on reasoning. In adversarial tests where models are misled with incorrect initial reasoning, biomedical models degrade sharply, while larger or RL-trained general models show more robustness. To address this, we train BioMed-R1 using fine-tuning and reinforcement learning on reasoning-heavy examples. It achieves the strongest performance among similarly sized models. Further gains may come from incorporating clinical case reports and training with adversarial and backtracking scenarios.
title Disentangling Reasoning and Knowledge in Medical Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.11462