Instance-Level Composed Image Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Psomas, Bill, Retsinas, George, Efthymiadis, Nikos, Filntisis, Panagiotis, Avrithis, Yannis, Maragos, Petros, Chum, Ondrej, Tolias, Giorgos
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914201359876096
author Psomas, Bill
Retsinas, George
Efthymiadis, Nikos
Filntisis, Panagiotis
Avrithis, Yannis
Maragos, Petros
Chum, Ondrej
Tolias, Giorgos
author_facet Psomas, Bill
Retsinas, George
Efthymiadis, Nikos
Filntisis, Panagiotis
Avrithis, Yannis
Maragos, Petros
Chum, Ondrej
Tolias, Giorgos
contents The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge-comparable to retrieval among more than 40M random distractors-through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25387
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Instance-Level Composed Image Retrieval
Psomas, Bill
Retsinas, George
Efthymiadis, Nikos
Filntisis, Panagiotis
Avrithis, Yannis
Maragos, Petros
Chum, Ondrej
Tolias, Giorgos
Computer Vision and Pattern Recognition
The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge-comparable to retrieval among more than 40M random distractors-through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/.
title Instance-Level Composed Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.25387