See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Ruinan, Huang, Gexin, Shen, Xinwei, Zhang, Qiong, Tan, Yan Shuo, Li, Xiaoxiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918348809306112
author Jin, Ruinan
Huang, Gexin
Shen, Xinwei
Zhang, Qiong
Tan, Yan Shuo
Li, Xiaoxiao
author_facet Jin, Ruinan
Huang, Gexin
Shen, Xinwei
Zhang, Qiong
Tan, Yan Shuo
Li, Xiaoxiao
contents Medical image diagnosis is challenging because many diseases resemble normal anatomy and exhibit substantial interpatient variability. Clinicians routinely rely on comparative diagnosis, such as referencing cross-patient healthy control images to identify subtle but clinically meaningful abnormalities. Although healthy reference images are abundant in practice, existing medical vision-language models (VLMs) primarily operate in a single-image or single-series setting and lack explicit mechanisms for comparative diagnosis. This work investigates whether incorporating clinically motivated comparison can enhance VLM performance. We show that providing VLMs with both a query image and a matched healthy reference image, accompanied by cross-patient comparative prompts, significantly improves diagnostic performance. This performance can be further augmented by lightweight supervised fine-tuning (SFT) on a small amount of data. At the same time, we evaluate multiple strategies for selecting reference images, including random sampling, demographic attribute matching, embedding-based retrieval, and cross-center selection, and find consistently strong performance across all settings. Finally, we investigate why comparative diagnosis is effective theoretically, and observe improved sample efficiency and tighter alignment between visual and textual representations. Our findings highlight the clinical relevance of comparison-based diagnosis, provide practical strategies for incorporating reference images into VLMs, and demonstrate improved performance across diverse medical imaging tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18140
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
Jin, Ruinan
Huang, Gexin
Shen, Xinwei
Zhang, Qiong
Tan, Yan Shuo
Li, Xiaoxiao
Computer Vision and Pattern Recognition
Medical image diagnosis is challenging because many diseases resemble normal anatomy and exhibit substantial interpatient variability. Clinicians routinely rely on comparative diagnosis, such as referencing cross-patient healthy control images to identify subtle but clinically meaningful abnormalities. Although healthy reference images are abundant in practice, existing medical vision-language models (VLMs) primarily operate in a single-image or single-series setting and lack explicit mechanisms for comparative diagnosis. This work investigates whether incorporating clinically motivated comparison can enhance VLM performance. We show that providing VLMs with both a query image and a matched healthy reference image, accompanied by cross-patient comparative prompts, significantly improves diagnostic performance. This performance can be further augmented by lightweight supervised fine-tuning (SFT) on a small amount of data. At the same time, we evaluate multiple strategies for selecting reference images, including random sampling, demographic attribute matching, embedding-based retrieval, and cross-center selection, and find consistently strong performance across all settings. Finally, we investigate why comparative diagnosis is effective theoretically, and observe improved sample efficiency and tighter alignment between visual and textual representations. Our findings highlight the clinical relevance of comparison-based diagnosis, provide practical strategies for incorporating reference images into VLMs, and demonstrate improved performance across diverse medical imaging tasks.
title See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.18140