Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Haoyu, Song, Yaoxian, Wang, Xuwu, Xiangru, Zhu, Li, Zhixu, Song, Wei, Li, Tiefeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911819958845440
author Liu, Haoyu
Song, Yaoxian
Wang, Xuwu
Xiangru, Zhu
Li, Zhixu
Song, Wei
Li, Tiefeng
author_facet Liu, Haoyu
Song, Yaoxian
Wang, Xuwu
Xiangru, Zhu
Li, Zhixu
Song, Wei
Li, Tiefeng
contents With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is needed to realize high-quality and efficient retrieval between different modalities. Existing text-image retrieval research is mostly based on general vision-language datasets (e.g. MS-COCO, Flickr30K), in which the query utterance is rigid and unnatural (i.e. verbosity and formality). To overcome the shortcoming, we construct a new Compact and Fragmented Query challenge dataset (named Flickr30K-CFQ) to model text-image retrieval task considering multiple query content and style, including compact and fine-grained entity-relation corpus. We propose a novel query-enhanced text-image retrieval method using prompt engineering based on LLM. Experiments show that our proposed Flickr30-CFQ reveals the insufficiency of existing vision-language datasets in realistic text-image tasks. Our LLM-based Query-enhanced method applied on different existing text-image retrieval models improves query understanding performance both on public dataset and our challenge set Flickr30-CFQ with over 0.9% and 2.4% respectively. Our project can be available anonymously in https://sites.google.com/view/Flickr30K-cfq.
format Preprint
id arxiv_https___arxiv_org_abs_2403_13317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
Liu, Haoyu
Song, Yaoxian
Wang, Xuwu
Xiangru, Zhu
Li, Zhixu
Song, Wei
Li, Tiefeng
Information Retrieval
With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is needed to realize high-quality and efficient retrieval between different modalities. Existing text-image retrieval research is mostly based on general vision-language datasets (e.g. MS-COCO, Flickr30K), in which the query utterance is rigid and unnatural (i.e. verbosity and formality). To overcome the shortcoming, we construct a new Compact and Fragmented Query challenge dataset (named Flickr30K-CFQ) to model text-image retrieval task considering multiple query content and style, including compact and fine-grained entity-relation corpus. We propose a novel query-enhanced text-image retrieval method using prompt engineering based on LLM. Experiments show that our proposed Flickr30-CFQ reveals the insufficiency of existing vision-language datasets in realistic text-image tasks. Our LLM-based Query-enhanced method applied on different existing text-image retrieval models improves query understanding performance both on public dataset and our challenge set Flickr30-CFQ with over 0.9% and 2.4% respectively. Our project can be available anonymously in https://sites.google.com/view/Flickr30K-cfq.
title Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2403.13317