Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hashmi, Anees Ur Rehman, Mahapatra, Dwarikanath, Yaqub, Mohammad
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914731582816256
author Hashmi, Anees Ur Rehman
Mahapatra, Dwarikanath
Yaqub, Mohammad
author_facet Hashmi, Anees Ur Rehman
Mahapatra, Dwarikanath
Yaqub, Mohammad
contents Explaining Deep Learning models is becoming increasingly important in the face of daily emerging multimodal models, particularly in safety-critical domains like medical imaging. However, the lack of detailed investigations into the performance of explainability methods on these models is widening the gap between their development and safe deployment. In this work, we analyze the performance of various explainable AI methods on a vision-language model, MedCLIP, to demystify its inner workings. We also provide a simple methodology to overcome the shortcomings of these methods. Our work offers a different new perspective on the explainability of a recent well-known VLM in the medical domain and our assessment method is generalizable to other current and possible future VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2403_18996
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models
Hashmi, Anees Ur Rehman
Mahapatra, Dwarikanath
Yaqub, Mohammad
Computer Vision and Pattern Recognition
Explaining Deep Learning models is becoming increasingly important in the face of daily emerging multimodal models, particularly in safety-critical domains like medical imaging. However, the lack of detailed investigations into the performance of explainability methods on these models is widening the gap between their development and safe deployment. In this work, we analyze the performance of various explainable AI methods on a vision-language model, MedCLIP, to demystify its inner workings. We also provide a simple methodology to overcome the shortcomings of these methods. Our work offers a different new perspective on the explainability of a recent well-known VLM in the medical domain and our assessment method is generalizable to other current and possible future VLMs.
title Envisioning MedCLIP: A Deep Dive into Explainability for Medical Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.18996