PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pham, Thang M., Chen, Peijie, Nguyen, Tin, Yoon, Seunghyun, Bui, Trung, Nguyen, Anh Totti
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910409347301376
author Pham, Thang M.
Chen, Peijie
Nguyen, Tin
Yoon, Seunghyun
Bui, Trung
Nguyen, Anh Totti
author_facet Pham, Thang M.
Chen, Peijie
Nguyen, Tin
Yoon, Seunghyun
Bui, Trung
Nguyen, Anh Totti
contents CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rarely appear on the Internet (e.g., scientific names of birds). For fine-grained classification, we propose PEEB - an explainable and editable classifier to (1) express the class name into a set of text descriptors that describe the visual parts of that class; and (2) match the embeddings of the detected parts to their textual descriptors in each class to compute a logit score for classification. In a zero-shot setting where the class names are unknown, PEEB outperforms CLIP by a huge margin (~10x in top-1 accuracy). Compared to part-based classifiers, PEEB is not only the state-of-the-art (SOTA) on the supervised-learning setting (88.80% and 92.20% accuracy on CUB-200 and Dogs-120, respectively) but also the first to enable users to edit the text descriptors to form a new classifier without any re-training. Compared to concept bottleneck models, PEEB is also the SOTA in both zero-shot and supervised-learning settings.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05297
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
Pham, Thang M.
Chen, Peijie
Nguyen, Tin
Yoon, Seunghyun
Bui, Trung
Nguyen, Anh Totti
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
CLIP-based classifiers rely on the prompt containing a {class name} that is known to the text encoder. Therefore, they perform poorly on new classes or the classes whose names rarely appear on the Internet (e.g., scientific names of birds). For fine-grained classification, we propose PEEB - an explainable and editable classifier to (1) express the class name into a set of text descriptors that describe the visual parts of that class; and (2) match the embeddings of the detected parts to their textual descriptors in each class to compute a logit score for classification. In a zero-shot setting where the class names are unknown, PEEB outperforms CLIP by a huge margin (~10x in top-1 accuracy). Compared to part-based classifiers, PEEB is not only the state-of-the-art (SOTA) on the supervised-learning setting (88.80% and 92.20% accuracy on CUB-200 and Dogs-120, respectively) but also the first to enable users to edit the text descriptors to form a new classifier without any re-training. Compared to concept bottleneck models, PEEB is also the SOTA in both zero-shot and supervised-learning settings.
title PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2403.05297