EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huzaifa, Muhammad, Kementchedjhieva, Yova
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911200028131328
author Huzaifa, Muhammad
Kementchedjhieva, Yova
author_facet Huzaifa, Muhammad
Kementchedjhieva, Yova
contents Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-world complexity. Pre-trained vision-language models tend to perform well with easy negatives but struggle with hard negatives--visually similar yet incorrect images--especially in open-domain scenarios. To address this, we introduce Episodic Few-Shot Adaptation (EFSA), a novel test-time framework that adapts pre-trained models dynamically to a query's domain by fine-tuning on top-k retrieved candidates and synthetic captions generated for them. EFSA improves performance across diverse domains while preserving generalization, as shown in evaluations on queries from eight highly distinct visual domains and an open-domain retrieval pool of over one million images. Our work highlights the potential of episodic few-shot adaptation to enhance robustness in the critical and understudied task of open-domain text-to-image retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2412_00139
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
Huzaifa, Muhammad
Kementchedjhieva, Yova
Computer Vision and Pattern Recognition
Text-to-image retrieval is a critical task for managing diverse visual content, but common benchmarks for the task rely on small, single-domain datasets that fail to capture real-world complexity. Pre-trained vision-language models tend to perform well with easy negatives but struggle with hard negatives--visually similar yet incorrect images--especially in open-domain scenarios. To address this, we introduce Episodic Few-Shot Adaptation (EFSA), a novel test-time framework that adapts pre-trained models dynamically to a query's domain by fine-tuning on top-k retrieved candidates and synthetic captions generated for them. EFSA improves performance across diverse domains while preserving generalization, as shown in evaluations on queries from eight highly distinct visual domains and an open-domain retrieval pool of over one million images. Our work highlights the potential of episodic few-shot adaptation to enhance robustness in the critical and understudied task of open-domain text-to-image retrieval.
title EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.00139