Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Shidara, Kie, Prem, Preethi, Kim, Jonathan, Podlasek, Anna, Liu, Feng, Alaa, Ahmed, Bernardo, Danilo
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909993252421632
author Shidara, Kie
Prem, Preethi
Kim, Jonathan
Podlasek, Anna
Liu, Feng
Alaa, Ahmed
Bernardo, Danilo
author_facet Shidara, Kie
Prem, Preethi
Kim, Jonathan
Podlasek, Anna
Liu, Feng
Alaa, Ahmed
Bernardo, Danilo
contents Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC. On questions most commonly missed by physicians, the top 5 performing models answered 55% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11866
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving
Shidara, Kie
Prem, Preethi
Kim, Jonathan
Podlasek, Anna
Liu, Feng
Alaa, Ahmed
Bernardo, Danilo
Computation and Language
Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC. On questions most commonly missed by physicians, the top 5 performing models answered 55% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.
title Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving
topic Computation and Language
url https://arxiv.org/abs/2601.11866