Nearest Neighbor Normalization Improves Multimodal Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chowdhury, Neil, Wang, Franklin, Shenoy, Sumedh, Kiela, Douwe, Schwettmann, Sarah, Thrush, Tristan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912098496282624
author Chowdhury, Neil
Wang, Franklin
Shenoy, Sumedh
Kiela, Douwe
Schwettmann, Sarah
Thrush, Tristan
author_facet Chowdhury, Neil
Wang, Franklin
Shenoy, Sumedh
Kiela, Douwe
Schwettmann, Sarah
Thrush, Tristan
contents Multimodal models leverage large-scale pre-training to achieve strong but still imperfect performance on tasks such as image captioning, visual question answering, and cross-modal retrieval. In this paper, we present a simple and efficient method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization (NNN). We show an improvement on retrieval metrics in both text retrieval and image retrieval for all of the contrastive models that we tested (CLIP, BLIP, ALBEF, SigLIP, BEiT) and for both of the datasets that we used (MS-COCO and Flickr30k). NNN requires a reference database, but does not require any training on this database, and can even increase the retrieval accuracy of a model after finetuning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_24114
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Nearest Neighbor Normalization Improves Multimodal Retrieval
Chowdhury, Neil
Wang, Franklin
Shenoy, Sumedh
Kiela, Douwe
Schwettmann, Sarah
Thrush, Tristan
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Multimodal models leverage large-scale pre-training to achieve strong but still imperfect performance on tasks such as image captioning, visual question answering, and cross-modal retrieval. In this paper, we present a simple and efficient method for correcting errors in trained contrastive image-text retrieval models with no additional training, called Nearest Neighbor Normalization (NNN). We show an improvement on retrieval metrics in both text retrieval and image retrieval for all of the contrastive models that we tested (CLIP, BLIP, ALBEF, SigLIP, BEiT) and for both of the datasets that we used (MS-COCO and Flickr30k). NNN requires a reference database, but does not require any training on this database, and can even increase the retrieval accuracy of a model after finetuning.
title Nearest Neighbor Normalization Improves Multimodal Retrieval
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.24114