Evidential Transformers for Improved Image Retrieval

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Dordevic, Danilo, Kumar, Suryansh
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916938282696704
author Dordevic, Danilo
Kumar, Suryansh
author_facet Dordevic, Danilo
Kumar, Suryansh
contents We introduce the Evidential Transformer, an uncertainty-driven transformer model for improved and robust image retrieval. In this paper, we make several contributions to content-based image retrieval (CBIR). We incorporate probabilistic methods into image retrieval, achieving robust and reliable results, with evidential classification surpassing traditional training based on multiclass classification as a baseline for deep metric learning. Furthermore, we improve the state-of-the-art retrieval results on several datasets by leveraging the Global Context Vision Transformer (GC ViT) architecture. Our experimental results consistently demonstrate the reliability of our approach, setting a new benchmark in CBIR in all test settings on the Stanford Online Products (SOP) and CUB-200-2011 datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01082
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evidential Transformers for Improved Image Retrieval
Dordevic, Danilo
Kumar, Suryansh
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
We introduce the Evidential Transformer, an uncertainty-driven transformer model for improved and robust image retrieval. In this paper, we make several contributions to content-based image retrieval (CBIR). We incorporate probabilistic methods into image retrieval, achieving robust and reliable results, with evidential classification surpassing traditional training based on multiclass classification as a baseline for deep metric learning. Furthermore, we improve the state-of-the-art retrieval results on several datasets by leveraging the Global Context Vision Transformer (GC ViT) architecture. Our experimental results consistently demonstrate the reliability of our approach, setting a new benchmark in CBIR in all test settings on the Stanford Online Products (SOP) and CUB-200-2011 datasets.
title Evidential Transformers for Improved Image Retrieval
topic Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2409.01082