Towards Concept-based Interpretability of Skin Lesion Diagnosis using Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patrício, Cristiano, Teixeira, Luís F., Neves, João C.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911789882540032
author Patrício, Cristiano
Teixeira, Luís F.
Neves, João C.
author_facet Patrício, Cristiano
Teixeira, Luís F.
Neves, João C.
contents Concept-based models naturally lend themselves to the development of inherently interpretable skin lesion diagnosis, as medical experts make decisions based on a set of visual patterns of the lesion. Nevertheless, the development of these models depends on the existence of concept-annotated datasets, whose availability is scarce due to the specialized knowledge and expertise required in the annotation process. In this work, we show that vision-language models can be used to alleviate the dependence on a large number of concept-annotated samples. In particular, we propose an embedding learning strategy to adapt CLIP to the downstream task of skin lesion classification using concept-based descriptions as textual embeddings. Our experiments reveal that vision-language models not only attain better accuracy when using concepts as textual embeddings, but also require a smaller number of concept-annotated samples to attain comparable performance to approaches specifically devised for automatic concept generation.
format Preprint
id arxiv_https___arxiv_org_abs_2311_14339
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Towards Concept-based Interpretability of Skin Lesion Diagnosis using Vision-Language Models
Patrício, Cristiano
Teixeira, Luís F.
Neves, João C.
Computer Vision and Pattern Recognition
Concept-based models naturally lend themselves to the development of inherently interpretable skin lesion diagnosis, as medical experts make decisions based on a set of visual patterns of the lesion. Nevertheless, the development of these models depends on the existence of concept-annotated datasets, whose availability is scarce due to the specialized knowledge and expertise required in the annotation process. In this work, we show that vision-language models can be used to alleviate the dependence on a large number of concept-annotated samples. In particular, we propose an embedding learning strategy to adapt CLIP to the downstream task of skin lesion classification using concept-based descriptions as textual embeddings. Our experiments reveal that vision-language models not only attain better accuracy when using concepts as textual embeddings, but also require a smaller number of concept-annotated samples to attain comparable performance to approaches specifically devised for automatic concept generation.
title Towards Concept-based Interpretability of Skin Lesion Diagnosis using Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.14339