Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moon, Jihyun, Hong, Charmgil
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911147043586048
author Moon, Jihyun
Hong, Charmgil
author_facet Moon, Jihyun
Hong, Charmgil
contents Accurate and early diagnosis of malignant melanoma is critical for improving patient outcomes. While convolutional neural networks (CNNs) have shown promise in dermoscopic image analysis, they often neglect clinical metadata and require extensive preprocessing. Vision-language models (VLMs) offer a multimodal alternative but struggle to capture clinical specificity when trained on general-domain data. To address this, we propose a retrieval-augmented VLM framework that incorporates semantically similar patient cases into the diagnostic prompt. Our method enables informed predictions without fine-tuning and significantly improves classification accuracy and error correction over conventional baselines. These results demonstrate that retrieval-augmented prompting provides a robust strategy for clinical decision support.
format Preprint
id arxiv_https___arxiv_org_abs_2509_08338
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
Moon, Jihyun
Hong, Charmgil
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Accurate and early diagnosis of malignant melanoma is critical for improving patient outcomes. While convolutional neural networks (CNNs) have shown promise in dermoscopic image analysis, they often neglect clinical metadata and require extensive preprocessing. Vision-language models (VLMs) offer a multimodal alternative but struggle to capture clinical specificity when trained on general-domain data. To address this, we propose a retrieval-augmented VLM framework that incorporates semantically similar patient cases into the diagnostic prompt. Our method enables informed predictions without fine-tuning and significantly improves classification accuracy and error correction over conventional baselines. These results demonstrate that retrieval-augmented prompting provides a robust strategy for clinical decision support.
title Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.08338