Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Duan, Yiqun, Ramasinghe, Sameera, Gould, Stephen, Thalaiyasingam, Ajanthan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918057570467840
author Duan, Yiqun
Ramasinghe, Sameera
Gould, Stephen
Thalaiyasingam, Ajanthan
author_facet Duan, Yiqun
Ramasinghe, Sameera
Gould, Stephen
Thalaiyasingam, Ajanthan
contents Composed Image Retrieval (CIR) is the task of retrieving images matching a reference image augmented with a text, where the text describes changes to the reference image in natural language. Traditionally, models designed for CIR have relied on triplet data containing a reference image, reformulation text, and a target image. However, curating such triplet data often necessitates human intervention, leading to prohibitive costs. This challenge has hindered the scalability of CIR model training even with the availability of abundant unlabeled data. With the recent advances in foundational models, we advocate a shift in the CIR training paradigm where human annotations can be efficiently replaced by large language models (LLMs). Specifically, we demonstrate the capability of large captioning and language models in efficiently generating data for CIR only relying on unannotated image collections. Additionally, we introduce an embedding reformulation architecture that effectively combines image and text modalities. Our model, named InstructCIR, outperforms state-of-the-art methods in zero-shot composed image retrieval on CIRR and FashionIQ datasets. Furthermore, we demonstrate that by increasing the amount of generated data, our zero-shot model gets closer to the performance of supervised baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2504_00812
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
Duan, Yiqun
Ramasinghe, Sameera
Gould, Stephen
Thalaiyasingam, Ajanthan
Computer Vision and Pattern Recognition
Multimedia
Composed Image Retrieval (CIR) is the task of retrieving images matching a reference image augmented with a text, where the text describes changes to the reference image in natural language. Traditionally, models designed for CIR have relied on triplet data containing a reference image, reformulation text, and a target image. However, curating such triplet data often necessitates human intervention, leading to prohibitive costs. This challenge has hindered the scalability of CIR model training even with the availability of abundant unlabeled data. With the recent advances in foundational models, we advocate a shift in the CIR training paradigm where human annotations can be efficiently replaced by large language models (LLMs). Specifically, we demonstrate the capability of large captioning and language models in efficiently generating data for CIR only relying on unannotated image collections. Additionally, we introduce an embedding reformulation architecture that effectively combines image and text modalities. Our model, named InstructCIR, outperforms state-of-the-art methods in zero-shot composed image retrieval on CIRR and FashionIQ datasets. Furthermore, we demonstrate that by increasing the amount of generated data, our zero-shot model gets closer to the performance of supervised baselines.
title Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2504.00812