Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Demić, Emil, Zajc, Luka Čehovin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916939746508800
author Demić, Emil
Zajc, Luka Čehovin
author_facet Demić, Emil
Zajc, Luka Čehovin
contents The goal of Scene-level Sketch-Based Image Retrieval is to retrieve natural images matching the overall semantics and spatial layout of a free-hand sketch. Unlike prior work focused on architectural augmentations of retrieval models, we emphasize the inherent ambiguity and noise present in real-world sketches. This insight motivates a training objective that is explicitly designed to be robust to sketch variability. We show that with an appropriate combination of pre-training, encoder architecture, and loss formulation, it is possible to achieve state-of-the-art performance without the introduction of additional complexity. Extensive experiments on a challenging FS-COCO and widely-used SketchyCOCO datasets confirm the effectiveness of our approach and underline the critical role of training design in cross-modal retrieval tasks, as well as the need to improve the evaluation scenarios of scene-level SBIR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval
Demić, Emil
Zajc, Luka Čehovin
Computer Vision and Pattern Recognition
The goal of Scene-level Sketch-Based Image Retrieval is to retrieve natural images matching the overall semantics and spatial layout of a free-hand sketch. Unlike prior work focused on architectural augmentations of retrieval models, we emphasize the inherent ambiguity and noise present in real-world sketches. This insight motivates a training objective that is explicitly designed to be robust to sketch variability. We show that with an appropriate combination of pre-training, encoder architecture, and loss formulation, it is possible to achieve state-of-the-art performance without the introduction of additional complexity. Extensive experiments on a challenging FS-COCO and widely-used SketchyCOCO datasets confirm the effectiveness of our approach and underline the critical role of training design in cross-modal retrieval tasks, as well as the need to improve the evaluation scenarios of scene-level SBIR.
title Back To The Drawing Board: Rethinking Scene-Level Sketch-Based Image Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.06566