Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Korbar, Bruno, Zisserman, Andrew
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912632522407936
author Korbar, Bruno
Zisserman, Andrew
author_facet Korbar, Bruno
Zisserman, Andrew
contents The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that object is doing or where it is. For example, to retrieve an image of "Fluffy the unicorn (specified by an image) on someone's head". To achieve this we design a mapping network that can "translate" from a local image embedding (of the object instance) to a text token, such that the combination of the token and a natural language query is suitable for CLIP style text encoding, and image retrieval. Generating a text token in this manner involves a simple training procedure, that only needs to be performed once for each object instance. We show that our approach of using a trainable mapping network, termed pi-map, together with frozen CLIP text and image encoders, improves the state of the art on two benchmarks designed to assess personalized retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05411
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
Korbar, Bruno
Zisserman, Andrew
Computer Vision and Pattern Recognition
The goal of this paper is to be able to retrieve images using a compound query that combines object instance information from an image, with a natural text description of what that object is doing or where it is. For example, to retrieve an image of "Fluffy the unicorn (specified by an image) on someone's head". To achieve this we design a mapping network that can "translate" from a local image embedding (of the object instance) to a text token, such that the combination of the token and a natural language query is suitable for CLIP style text encoding, and image retrieval. Generating a text token in this manner involves a simple training procedure, that only needs to be performed once for each object instance. We show that our approach of using a trainable mapping network, termed pi-map, together with frozen CLIP text and image encoders, improves the state of the art on two benchmarks designed to assess personalized retrieval.
title Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.05411