Zero-shot Composed Text-Image Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yikun, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng, Xie, Weidi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917605155012608
author Liu, Yikun
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
author_facet Liu, Yikun
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
contents In this paper, we consider the problem of composed image retrieval (CIR), it aims to train a model that can fuse multi-modal information, e.g., text and images, to accurately retrieve images that match the query, extending the user's expression ability. We make the following contributions: (i) we initiate a scalable pipeline to automatically construct datasets for training CIR model, by simply exploiting a large-scale dataset of image-text pairs, e.g., a subset of LAION-5B; (ii) we introduce a transformer-based adaptive aggregation model, TransAgg, which employs a simple yet efficient fusion mechanism, to adaptively combine information from diverse modalities; (iii) we conduct extensive ablation studies to investigate the usefulness of our proposed data construction procedure, and the effectiveness of core components in TransAgg; (iv) when evaluating on the publicly available benckmarks under the zero-shot scenario, i.e., training on the automatically constructed datasets, then directly conduct inference on target downstream datasets, e.g., CIRR and FashionIQ, our proposed approach either performs on par with or significantly outperforms the existing state-of-the-art (SOTA) models. Project page: https://code-kunkun.github.io/ZS-CIR/
format Preprint
id arxiv_https___arxiv_org_abs_2306_07272
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-shot Composed Text-Image Retrieval
Liu, Yikun
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Xie, Weidi
Computer Vision and Pattern Recognition
In this paper, we consider the problem of composed image retrieval (CIR), it aims to train a model that can fuse multi-modal information, e.g., text and images, to accurately retrieve images that match the query, extending the user's expression ability. We make the following contributions: (i) we initiate a scalable pipeline to automatically construct datasets for training CIR model, by simply exploiting a large-scale dataset of image-text pairs, e.g., a subset of LAION-5B; (ii) we introduce a transformer-based adaptive aggregation model, TransAgg, which employs a simple yet efficient fusion mechanism, to adaptively combine information from diverse modalities; (iii) we conduct extensive ablation studies to investigate the usefulness of our proposed data construction procedure, and the effectiveness of core components in TransAgg; (iv) when evaluating on the publicly available benckmarks under the zero-shot scenario, i.e., training on the automatically constructed datasets, then directly conduct inference on target downstream datasets, e.g., CIRR and FashionIQ, our proposed approach either performs on par with or significantly outperforms the existing state-of-the-art (SOTA) models. Project page: https://code-kunkun.github.io/ZS-CIR/
title Zero-shot Composed Text-Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.07272