A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916286589566976 |
|---|---|
| author | Paul, Dipanjyoti Chowdhury, Arpita Xiong, Xinqi Chang, Feng-Ju Carlyn, David Stevens, Samuel Provost, Kaiya L. Karpatne, Anuj Carstens, Bryan Rubenstein, Daniel Stewart, Charles Berger-Wolf, Tanya Su, Yu Chao, Wei-Lun |
| author_facet | Paul, Dipanjyoti Chowdhury, Arpita Xiong, Xinqi Chang, Feng-Ju Carlyn, David Stevens, Samuel Provost, Kaiya L. Karpatne, Anuj Carstens, Bryan Rubenstein, Daniel Stewart, Charles Berger-Wolf, Tanya Su, Yu Chao, Wei-Lun |
| contents | We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn "class-specific" queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via "multi-head" cross-attention, INTR could identify different "attributes" of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_04157 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis Paul, Dipanjyoti Chowdhury, Arpita Xiong, Xinqi Chang, Feng-Ju Carlyn, David Stevens, Samuel Provost, Kaiya L. Karpatne, Anuj Carstens, Bryan Rubenstein, Daniel Stewart, Charles Berger-Wolf, Tanya Su, Yu Chao, Wei-Lun Computer Vision and Pattern Recognition Artificial Intelligence We present a novel usage of Transformers to make image classification interpretable. Unlike mainstream classifiers that wait until the last fully connected layer to incorporate class information to make predictions, we investigate a proactive approach, asking each class to search for itself in an image. We realize this idea via a Transformer encoder-decoder inspired by DEtection TRansformer (DETR). We learn "class-specific" queries (one for each class) as input to the decoder, enabling each class to localize its patterns in an image via cross-attention. We name our approach INterpretable TRansformer (INTR), which is fairly easy to implement and exhibits several compelling properties. We show that INTR intrinsically encourages each class to attend distinctively; the cross-attention weights thus provide a faithful interpretation of the prediction. Interestingly, via "multi-head" cross-attention, INTR could identify different "attributes" of a class, making it particularly suitable for fine-grained classification and analysis, which we demonstrate on eight datasets. Our code and pre-trained models are publicly accessible at the Imageomics Institute GitHub site: https://github.com/Imageomics/INTR. |
| title | A Simple Interpretable Transformer for Fine-Grained Image Classification and Analysis |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2311.04157 |