DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yilmaz, Abdurrahim, Erdem, Ozan, Gokyayla, Ece, Acar, Ayda, Dagtas, Burc Bugra, Erdil, Dilara Ilhan, Gencoglan, Gulsum, Temelkuran, Burak
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911387293319168
author Yilmaz, Abdurrahim
Erdem, Ozan
Gokyayla, Ece
Acar, Ayda
Dagtas, Burc Bugra
Erdil, Dilara Ilhan
Gencoglan, Gulsum
Temelkuran, Burak
author_facet Yilmaz, Abdurrahim
Erdem, Ozan
Gokyayla, Ece
Acar, Ayda
Dagtas, Burc Bugra
Erdil, Dilara Ilhan
Gencoglan, Gulsum
Temelkuran, Burak
contents Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.
format Preprint
id arxiv_https___arxiv_org_abs_2601_14084
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
Yilmaz, Abdurrahim
Erdem, Ozan
Gokyayla, Ece
Acar, Ayda
Dagtas, Burc Bugra
Erdil, Dilara Ilhan
Gencoglan, Gulsum
Temelkuran, Burak
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Vision-language models (VLMs) are increasingly important in medical applications; however, their evaluation in dermatology remains limited by datasets that focus primarily on image-level classification tasks such as lesion recognition. While valuable for recognition, such datasets cannot assess the full visual understanding, language grounding, and clinical reasoning capabilities of multimodal models. Visual question answering (VQA) benchmarks are required to evaluate how models interpret dermatological images, reason over fine-grained morphology, and generate clinically meaningful descriptions. We introduce DermaBench, a clinician-annotated dermatology VQA benchmark built on the Diverse Dermatology Images (DDI) dataset. DermaBench comprises 656 clinical images from 570 unique patients spanning Fitzpatrick skin types I-VI. Using a hierarchical annotation schema with 22 main questions (single-choice, multi-choice, and open-ended), expert dermatologists annotated each image for diagnosis, anatomic site, lesion morphology, distribution, surface features, color, and image quality, together with open-ended narrative descriptions and summaries, yielding approximately 14.474 VQA-style annotations. DermaBench is released as a metadata-only dataset to respect upstream licensing and is publicly available at Harvard Dataverse.
title DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.14084