Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Huaying, Yanagi, Rintaro, Togo, Ren, Ogawa, Takahiro, Haseyama, Miki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909232372121600
author Zhang, Huaying
Yanagi, Rintaro
Togo, Ren
Ogawa, Takahiro
Haseyama, Miki
author_facet Zhang, Huaying
Yanagi, Rintaro
Togo, Ren
Ogawa, Takahiro
Haseyama, Miki
contents This paper proposes a novel zero-shot composed image retrieval (CIR) method considering the query-target relationship by masked image-text pairs. The objective of CIR is to retrieve the target image using a query image and a query text. Existing methods use a textual inversion network to convert the query image into a pseudo word to compose the image and text and use a pre-trained visual-language model to realize the retrieval. However, they do not consider the query-target relationship to train the textual inversion network to acquire information for retrieval. In this paper, we propose a novel zero-shot CIR method that is trained end-to-end using masked image-text pairs. By exploiting the abundant image-text pairs that are convenient to obtain with a masking strategy for learning the query-target relationship, it is expected that accurate zero-shot CIR using a retrieval-focused textual inversion network can be realized. Experimental results show the effectiveness of the proposed method.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18836
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
Zhang, Huaying
Yanagi, Rintaro
Togo, Ren
Ogawa, Takahiro
Haseyama, Miki
Computer Vision and Pattern Recognition
Information Retrieval
This paper proposes a novel zero-shot composed image retrieval (CIR) method considering the query-target relationship by masked image-text pairs. The objective of CIR is to retrieve the target image using a query image and a query text. Existing methods use a textual inversion network to convert the query image into a pseudo word to compose the image and text and use a pre-trained visual-language model to realize the retrieval. However, they do not consider the query-target relationship to train the textual inversion network to acquire information for retrieval. In this paper, we propose a novel zero-shot CIR method that is trained end-to-end using masked image-text pairs. By exploiting the abundant image-text pairs that are convenient to obtain with a masking strategy for learning the query-target relationship, it is expected that accurate zero-shot CIR using a retrieval-focused textual inversion network can be realized. Experimental results show the effectiveness of the proposed method.
title Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
topic Computer Vision and Pattern Recognition
Information Retrieval
url https://arxiv.org/abs/2406.18836