DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yilmaz, Abdurrahim, Yuceyalcin, Furkan, Gokyayla, Ece, Choi, Donghee, Erdem, Ozan, Demircali, Ali Anil, Varol, Rahmetullah, Kirabali, Ufuk Gorkem, Gencoglan, Gulsum, Posma, Joram M., Temelkuran, Burak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910856749514752
author Yilmaz, Abdurrahim
Yuceyalcin, Furkan
Gokyayla, Ece
Choi, Donghee
Erdem, Ozan
Demircali, Ali Anil
Varol, Rahmetullah
Kirabali, Ufuk Gorkem
Gencoglan, Gulsum
Posma, Joram M.
Temelkuran, Burak
author_facet Yilmaz, Abdurrahim
Yuceyalcin, Furkan
Gokyayla, Ece
Choi, Donghee
Erdem, Ozan
Demircali, Ali Anil
Varol, Rahmetullah
Kirabali, Ufuk Gorkem
Gencoglan, Gulsum
Posma, Joram M.
Temelkuran, Burak
contents A major barrier to developing vision large language models (LLMs) in dermatology is the lack of large image--text pairs dataset. We introduce DermaSynth, a dataset comprising of 92,020 synthetic image--text pairs curated from 45,205 images (13,568 clinical and 35,561 dermatoscopic) for dermatology-related clinical tasks. Leveraging state-of-the-art LLMs, using Gemini 2.0, we used clinically related prompts and self-instruct method to generate diverse and rich synthetic texts. Metadata of the datasets were incorporated into the input prompts by targeting to reduce potential hallucinations. The resulting dataset builds upon open access dermatological image repositories (DERM12345, BCN20000, PAD-UFES-20, SCIN, and HIBA) that have permissive CC-BY-4.0 licenses. We also fine-tuned a preliminary Llama-3.2-11B-Vision-Instruct model, DermatoLlama 1.0, on 5,000 samples. We anticipate this dataset to support and accelerate AI research in dermatology. Data and code underlying this work are accessible at https://github.com/abdurrahimyilmaz/DermaSynth.
format Preprint
id arxiv_https___arxiv_org_abs_2502_00196
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets
Yilmaz, Abdurrahim
Yuceyalcin, Furkan
Gokyayla, Ece
Choi, Donghee
Erdem, Ozan
Demircali, Ali Anil
Varol, Rahmetullah
Kirabali, Ufuk Gorkem
Gencoglan, Gulsum
Posma, Joram M.
Temelkuran, Burak
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
A major barrier to developing vision large language models (LLMs) in dermatology is the lack of large image--text pairs dataset. We introduce DermaSynth, a dataset comprising of 92,020 synthetic image--text pairs curated from 45,205 images (13,568 clinical and 35,561 dermatoscopic) for dermatology-related clinical tasks. Leveraging state-of-the-art LLMs, using Gemini 2.0, we used clinically related prompts and self-instruct method to generate diverse and rich synthetic texts. Metadata of the datasets were incorporated into the input prompts by targeting to reduce potential hallucinations. The resulting dataset builds upon open access dermatological image repositories (DERM12345, BCN20000, PAD-UFES-20, SCIN, and HIBA) that have permissive CC-BY-4.0 licenses. We also fine-tuned a preliminary Llama-3.2-11B-Vision-Instruct model, DermatoLlama 1.0, on 5,000 samples. We anticipate this dataset to support and accelerate AI research in dermatology. Data and code underlying this work are accessible at https://github.com/abdurrahimyilmaz/DermaSynth.
title DermaSynth: Rich Synthetic Image-Text Pairs Using Open Access Dermatology Datasets
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.00196