Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909232372121600 |
|---|---|
| author | Zhang, Huaying Yanagi, Rintaro Togo, Ren Ogawa, Takahiro Haseyama, Miki |
| author_facet | Zhang, Huaying Yanagi, Rintaro Togo, Ren Ogawa, Takahiro Haseyama, Miki |
| contents | This paper proposes a novel zero-shot composed image retrieval (CIR) method considering the query-target relationship by masked image-text pairs. The objective of CIR is to retrieve the target image using a query image and a query text. Existing methods use a textual inversion network to convert the query image into a pseudo word to compose the image and text and use a pre-trained visual-language model to realize the retrieval. However, they do not consider the query-target relationship to train the textual inversion network to acquire information for retrieval. In this paper, we propose a novel zero-shot CIR method that is trained end-to-end using masked image-text pairs. By exploiting the abundant image-text pairs that are convenient to obtain with a masking strategy for learning the query-target relationship, it is expected that accurate zero-shot CIR using a retrieval-focused textual inversion network can be realized. Experimental results show the effectiveness of the proposed method. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2406_18836 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs Zhang, Huaying Yanagi, Rintaro Togo, Ren Ogawa, Takahiro Haseyama, Miki Computer Vision and Pattern Recognition Information Retrieval This paper proposes a novel zero-shot composed image retrieval (CIR) method considering the query-target relationship by masked image-text pairs. The objective of CIR is to retrieve the target image using a query image and a query text. Existing methods use a textual inversion network to convert the query image into a pseudo word to compose the image and text and use a pre-trained visual-language model to realize the retrieval. However, they do not consider the query-target relationship to train the textual inversion network to acquire information for retrieval. In this paper, we propose a novel zero-shot CIR method that is trained end-to-end using masked image-text pairs. By exploiting the abundant image-text pairs that are convenient to obtain with a masking strategy for learning the query-target relationship, it is expected that accurate zero-shot CIR using a retrieval-focused textual inversion network can be realized. Experimental results show the effectiveness of the proposed method. |
| title | Zero-shot Composed Image Retrieval Considering Query-target Relationship Leveraging Masked Image-text Pairs |
| topic | Computer Vision and Pattern Recognition Information Retrieval |
| url | https://arxiv.org/abs/2406.18836 |