Medical thinking with multiple images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Zonghai, Wang, Benlu, Zhang, Yifan, Wang, Junda, Xia, Iris, Tang, Zhipeng, Han, Shuo, Ouyang, Feiyun, Yang, Zhichao, Cohan, Arman, Yu, Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918479061319680
author Yao, Zonghai
Wang, Benlu
Zhang, Yifan
Wang, Junda
Xia, Iris
Tang, Zhipeng
Han, Shuo
Ouyang, Feiyun
Yang, Zhichao
Cohan, Arman
Yu, Hong
author_facet Yao, Zonghai
Wang, Benlu
Zhang, Yifan
Wang, Junda
Xia, Iris
Tang, Zhipeng
Han, Shuo
Ouyang, Feiyun
Yang, Zhichao
Cohan, Arman
Yu, Hong
contents Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert-annotated benchmark for thinking with multiple images, where models must interpret each image, combine cross-view evidence, and answer diagnostic questions with intermediate supervision and step-level evaluation. The dataset contains 8,067 cases, including 720 test cases, with an average of 6.62 images per case, substantially denser than prior work, whose expert-level benchmarks use at most 1.43 images per case. On the test set, the best closed-source models, Claude-4.6-Opus, Gemini-3-Pro, and GPT-5.2-xhigh, reach only 57.2%, 55.3%, and 54.9% accuracy, while GPT-5-mini and GPT-5-nano reach 39.7% and 30.8%. Strong open-source models lag behind, led by Qwen3.5-397B-A17B at 52.2% and Qwen3.5-27B at 50.6%. Further analysis identifies grounded multi-image reasoning as the main bottleneck: models often fail to extract, align, and compose evidence across views before higher-level inference can help. Providing expert single-image cues and cross-image summaries improves performance, whereas replacing them with self-generated intermediates reduces accuracy. Step-level analysis shows that over 70% of errors arise from image reading and cross-view integration. Scaling results further show that additional inference-time computation helps only when visual grounding is already reliable; when early evidence extraction is weak, longer reasoning yields limited or unstable gains and can amplify misread cues. These results suggest that the key challenge is not reasoning length alone, but reliable mechanisms for grounding, aligning, and composing distributed evidence across real-world multimodal clinical inputs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16506
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Medical thinking with multiple images
Yao, Zonghai
Wang, Benlu
Zhang, Yifan
Wang, Junda
Xia, Iris
Tang, Zhipeng
Han, Shuo
Ouyang, Feiyun
Yang, Zhichao
Cohan, Arman
Yu, Hong
Computer Vision and Pattern Recognition
Computation and Language
Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a single view. We introduce MedThinkVQA, an expert-annotated benchmark for thinking with multiple images, where models must interpret each image, combine cross-view evidence, and answer diagnostic questions with intermediate supervision and step-level evaluation. The dataset contains 8,067 cases, including 720 test cases, with an average of 6.62 images per case, substantially denser than prior work, whose expert-level benchmarks use at most 1.43 images per case. On the test set, the best closed-source models, Claude-4.6-Opus, Gemini-3-Pro, and GPT-5.2-xhigh, reach only 57.2%, 55.3%, and 54.9% accuracy, while GPT-5-mini and GPT-5-nano reach 39.7% and 30.8%. Strong open-source models lag behind, led by Qwen3.5-397B-A17B at 52.2% and Qwen3.5-27B at 50.6%. Further analysis identifies grounded multi-image reasoning as the main bottleneck: models often fail to extract, align, and compose evidence across views before higher-level inference can help. Providing expert single-image cues and cross-image summaries improves performance, whereas replacing them with self-generated intermediates reduces accuracy. Step-level analysis shows that over 70% of errors arise from image reading and cross-view integration. Scaling results further show that additional inference-time computation helps only when visual grounding is already reliable; when early evidence extraction is weak, longer reasoning yields limited or unstable gains and can amplify misread cues. These results suggest that the key challenge is not reasoning length alone, but reliable mechanisms for grounding, aligning, and composing distributed evidence across real-world multimodal clinical inputs.
title Medical thinking with multiple images
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2604.16506