Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911965592420352 |
|---|---|
| author | Zhao, Ziwei Leake, David Ye, Xiaomeng Crandall, David |
| author_facet | Zhao, Ziwei Leake, David Ye, Xiaomeng Crandall, David |
| contents | This short paper presents preliminary research on the Case-Enhanced Vision Transformer (CEViT), a similarity measurement method aimed at improving the explainability of similarity assessments for image data. Initial experimental results suggest that integrating CEViT into k-Nearest Neighbor (k-NN) classification yields classification accuracy comparable to state-of-the-art computer vision models, while adding capabilities for illustrating differences between classes. CEViT explanations can be influenced by prior cases, to illustrate aspects of similarity relevant to those cases. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_16981 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric Zhao, Ziwei Leake, David Ye, Xiaomeng Crandall, David Computer Vision and Pattern Recognition Artificial Intelligence This short paper presents preliminary research on the Case-Enhanced Vision Transformer (CEViT), a similarity measurement method aimed at improving the explainability of similarity assessments for image data. Initial experimental results suggest that integrating CEViT into k-Nearest Neighbor (k-NN) classification yields classification accuracy comparable to state-of-the-art computer vision models, while adding capabilities for illustrating differences between classes. CEViT explanations can be influenced by prior cases, to illustrate aspects of similarity relevant to those cases. |
| title | Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2407.16981 |