Data-Efficient Generalization for Zero-shot Composed Image Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zining, Zhao, Zhicheng, Su, Fei, Zhang, Xiaoqin, Lu, Shijian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913723573075968
author Chen, Zining
Zhao, Zhicheng
Su, Fei
Zhang, Xiaoqin
Lu, Shijian
author_facet Chen, Zining
Zhao, Zhicheng
Su, Fei
Zhang, Xiaoqin
Lu, Shijian
contents Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text description without requiring in-distribution triplets for training. One prevalent approach follows the vision-language pretraining paradigm that employs a mapping network to transfer the image embedding to a pseudo-word token in the text embedding space. However, this approach tends to impede network generalization due to modality discrepancy and distribution shift between training and inference. To this end, we propose a Data-efficient Generalization (DeG) framework, including two novel designs, namely, Textual Supplement (TS) module and Semantic-Set (S-Set). The TS module exploits compositional textual semantics during training, enhancing the pseudo-word token with more linguistic semantics and thus mitigating the modality discrepancy effectively. The S-Set exploits the zero-shot capability of pretrained Vision-Language Models (VLMs), alleviating the distribution shift and mitigating the overfitting issue from the redundancy of the large-scale image-text data. Extensive experiments over four ZS-CIR benchmarks show that DeG outperforms the state-of-the-art (SOTA) methods with much less training data, and saves substantial training and inference time for practical usage.
format Preprint
id arxiv_https___arxiv_org_abs_2503_05204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data-Efficient Generalization for Zero-shot Composed Image Retrieval
Chen, Zining
Zhao, Zhicheng
Su, Fei
Zhang, Xiaoqin
Lu, Shijian
Computer Vision and Pattern Recognition
Zero-shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text description without requiring in-distribution triplets for training. One prevalent approach follows the vision-language pretraining paradigm that employs a mapping network to transfer the image embedding to a pseudo-word token in the text embedding space. However, this approach tends to impede network generalization due to modality discrepancy and distribution shift between training and inference. To this end, we propose a Data-efficient Generalization (DeG) framework, including two novel designs, namely, Textual Supplement (TS) module and Semantic-Set (S-Set). The TS module exploits compositional textual semantics during training, enhancing the pseudo-word token with more linguistic semantics and thus mitigating the modality discrepancy effectively. The S-Set exploits the zero-shot capability of pretrained Vision-Language Models (VLMs), alleviating the distribution shift and mitigating the overfitting issue from the redundancy of the large-scale image-text data. Extensive experiments over four ZS-CIR benchmarks show that DeG outperforms the state-of-the-art (SOTA) methods with much less training data, and saves substantial training and inference time for practical usage.
title Data-Efficient Generalization for Zero-shot Composed Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.05204