TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Qu, Leigang, Li, Haochuan, Wang, Tan, Wang, Wenjie, Li, Yongqi, Nie, Liqiang, Chua, Tat-Seng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910890610130944
author Qu, Leigang
Li, Haochuan
Wang, Tan
Wang, Wenjie
Li, Yongqi
Nie, Liqiang
Chua, Tat-Seng
author_facet Qu, Leigang
Li, Haochuan
Wang, Tan
Wang, Wenjie
Li, Yongqi
Nie, Liqiang
Chua, Tat-Seng
contents How humans can effectively and efficiently acquire images has always been a perennial question. A classic solution is text-to-image retrieval from an existing database; however, the limited database typically lacks creativity. By contrast, recent breakthroughs in text-to-image generation have made it possible to produce attractive and counterfactual visual content, but it faces challenges in synthesizing knowledge-intensive images. In this work, we rethink the relationship between text-to-image generation and retrieval, proposing a unified framework for both tasks with one single Large Multimodal Model (LMM). Specifically, we first explore the intrinsic discriminative abilities of LMMs and introduce an efficient generative retrieval method for text-to-image retrieval in a training-free manner. Subsequently, we unify generation and retrieval autoregressively and propose an autonomous decision mechanism to choose the best-matched one between generated and retrieved images as the response to the text prompt. To standardize the evaluation of unified text-to-image generation and retrieval, we construct TIGeR-Bench, a benchmark spanning both creative and knowledge-intensive domains. Extensive experiments on TIGeR-Bench and two retrieval benchmarks, i.e., Flickr30K and MS-COCO, demonstrate the superiority of our proposed framework.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
Qu, Leigang
Li, Haochuan
Wang, Tan
Wang, Wenjie
Li, Yongqi
Nie, Liqiang
Chua, Tat-Seng
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
How humans can effectively and efficiently acquire images has always been a perennial question. A classic solution is text-to-image retrieval from an existing database; however, the limited database typically lacks creativity. By contrast, recent breakthroughs in text-to-image generation have made it possible to produce attractive and counterfactual visual content, but it faces challenges in synthesizing knowledge-intensive images. In this work, we rethink the relationship between text-to-image generation and retrieval, proposing a unified framework for both tasks with one single Large Multimodal Model (LMM). Specifically, we first explore the intrinsic discriminative abilities of LMMs and introduce an efficient generative retrieval method for text-to-image retrieval in a training-free manner. Subsequently, we unify generation and retrieval autoregressively and propose an autonomous decision mechanism to choose the best-matched one between generated and retrieved images as the response to the text prompt. To standardize the evaluation of unified text-to-image generation and retrieval, we construct TIGeR-Bench, a benchmark spanning both creative and knowledge-intensive domains. Extensive experiments on TIGeR-Bench and two retrieval benchmarks, i.e., Flickr30K and MS-COCO, demonstrate the superiority of our proposed framework.
title TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2406.05814