Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zheyuan, Sun, Weixuan, Teney, Damien, Gould, Stephen
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914654878433280
author Liu, Zheyuan
Sun, Weixuan
Teney, Damien
Gould, Stephen
author_facet Liu, Zheyuan
Sun, Weixuan
Teney, Damien
Gould, Stephen
contents Composed image retrieval aims to find an image that best matches a given multi-modal user query consisting of a reference image and text pair. Existing methods commonly pre-compute image embeddings over the entire corpus and compare these to a reference image embedding modified by the query text at test time. Such a pipeline is very efficient at test time since fast vector distances can be used to evaluate candidates, but modifying the reference image embedding guided only by a short textual description can be difficult, especially independent of potential candidates. An alternative approach is to allow interactions between the query and every possible candidate, i.e., reference-text-candidate triplets, and pick the best from the entire set. Though this approach is more discriminative, for large-scale datasets the computational cost is prohibitive since pre-computation of candidate embeddings is no longer possible. We propose to combine the merits of both schemes using a two-stage model. Our first stage adopts the conventional vector distancing metric and performs a fast pruning among candidates. Meanwhile, our second stage employs a dual-encoder architecture, which effectively attends to the input triplet of reference-text-candidate and re-ranks the candidates. Both stages utilize a vision-and-language pre-trained network, which has proven beneficial for various downstream tasks. Our method consistently outperforms state-of-the-art approaches on standard benchmarks for the task. Our implementation is available at https://github.com/Cuberick-Orion/Candidate-Reranking-CIR.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16304
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
Liu, Zheyuan
Sun, Weixuan
Teney, Damien
Gould, Stephen
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
Composed image retrieval aims to find an image that best matches a given multi-modal user query consisting of a reference image and text pair. Existing methods commonly pre-compute image embeddings over the entire corpus and compare these to a reference image embedding modified by the query text at test time. Such a pipeline is very efficient at test time since fast vector distances can be used to evaluate candidates, but modifying the reference image embedding guided only by a short textual description can be difficult, especially independent of potential candidates. An alternative approach is to allow interactions between the query and every possible candidate, i.e., reference-text-candidate triplets, and pick the best from the entire set. Though this approach is more discriminative, for large-scale datasets the computational cost is prohibitive since pre-computation of candidate embeddings is no longer possible. We propose to combine the merits of both schemes using a two-stage model. Our first stage adopts the conventional vector distancing metric and performs a fast pruning among candidates. Meanwhile, our second stage employs a dual-encoder architecture, which effectively attends to the input triplet of reference-text-candidate and re-ranks the candidates. Both stages utilize a vision-and-language pre-trained network, which has proven beneficial for various downstream tasks. Our method consistently outperforms state-of-the-art approaches on standard benchmarks for the task. Our implementation is available at https://github.com/Cuberick-Orion/Candidate-Reranking-CIR.
title Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
topic Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2305.16304