A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paul, Dipanjyoti, Chowdhury, Arpita, Xiong, Xinqi, Chang, Feng-Ju, Carlyn, David, Stevens, Samuel, Provost, Kaiya L., Karpatne, Anuj, Carstens, Bryan, Rubenstein, Daniel, Stewart, Charles, Berger-Wolf, Tanya, Su, Yu, Chao, Wei-Lun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916286589566976
author Paul, Dipanjyoti
Chowdhury, Arpita
Xiong, Xinqi
Chang, Feng-Ju
Carlyn, David
Stevens, Samuel
Provost, Kaiya L.
Karpatne, Anuj
Carstens, Bryan
Rubenstein, Daniel
Stewart, Charles
Berger-Wolf, Tanya
Su, Yu
Chao, Wei-Lun
author_facet Paul, Dipanjyoti
Chowdhury, Arpita
Xiong, Xinqi
Chang, Feng-Ju
Carlyn, David
Stevens, Samuel
Provost, Kaiya L.
Karpatne, Anuj
Carstens, Bryan
Rubenstein, Daniel
Stewart, Charles
Berger-Wolf, Tanya
Su, Yu
Chao, Wei-Lun
contents We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn "class-specific" queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via "multi-head" cross-attention, INTR could identify different "attributes" of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR.
format Preprint
id arxiv_https___arxiv_org_abs_2311_04157
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
Paul, Dipanjyoti
Chowdhury, Arpita
Xiong, Xinqi
Chang, Feng-Ju
Carlyn, David
Stevens, Samuel
Provost, Kaiya L.
Karpatne, Anuj
Carstens, Bryan
Rubenstein, Daniel
Stewart, Charles
Berger-Wolf, Tanya
Su, Yu
Chao, Wei-Lun
Computer Vision and Pattern Recognition
Artificial Intelligence
We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn "class-specific" queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via "multi-head" cross-attention, INTR could identify different "attributes" of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR.
title A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2311.04157