ECOR: Explainable CLIP for Object Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rasekh, Ali, Ranjbar, Sepehr Kazemi, Heidari, Milad, Nejdl, Wolfgang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913322199154688
author Rasekh, Ali
Ranjbar, Sepehr Kazemi
Heidari, Milad
Nejdl, Wolfgang
author_facet Rasekh, Ali
Ranjbar, Sepehr Kazemi
Heidari, Milad
Nejdl, Wolfgang
contents Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. Their open vocabulary feature enhances their value. However, their black-box nature and lack of explainability in predictions make them less trustworthy in critical domains. Recently, some work has been done to force VLMs to provide reasonable rationales for object recognition, but this often comes at the expense of classification accuracy. In this paper, we first propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluations of different datasets, our method demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. The code will be made available online upon publication.
format Preprint
id arxiv_https___arxiv_org_abs_2404_12839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ECOR: Explainable CLIP for Object Recognition
Rasekh, Ali
Ranjbar, Sepehr Kazemi
Heidari, Milad
Nejdl, Wolfgang
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. Their open vocabulary feature enhances their value. However, their black-box nature and lack of explainability in predictions make them less trustworthy in critical domains. Recently, some work has been done to force VLMs to provide reasonable rationales for object recognition, but this often comes at the expense of classification accuracy. In this paper, we first propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluations of different datasets, our method demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. The code will be made available online upon publication.
title ECOR: Explainable CLIP for Object Recognition
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.12839