SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chaidos, Nikolaos, Dimitriou, Angeliki, Lymperaiou, Maria, Stamou, Giorgos
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909619524206592
author Chaidos, Nikolaos
Dimitriou, Angeliki
Lymperaiou, Maria
Stamou, Giorgos
author_facet Chaidos, Nikolaos
Dimitriou, Angeliki
Lymperaiou, Maria
Stamou, Giorgos
contents Despite the dominance of convolutional and transformer-based architectures in image-to-image retrieval, these models are prone to biases arising from low-level visual features, such as color. Recognizing the lack of semantic understanding as a key limitation, we propose a novel scene graph-based retrieval framework that emphasizes semantic content over superficial image characteristics. Prior approaches to scene graph retrieval predominantly rely on supervised Graph Neural Networks (GNNs), which require ground truth graph pairs driven from image captions. However, the inconsistency of caption-based supervision stemming from variable text encodings undermine retrieval reliability. To address these, we present SCENIR, a Graph Autoencoder-based unsupervised retrieval framework, which eliminates the dependence on labeled training data. Our model demonstrates superior performance across metrics and runtime efficiency, outperforming existing vision-based, multimodal, and supervised GNN approaches. We further advocate for Graph Edit Distance (GED) as a deterministic and robust ground truth measure for scene graph similarity, replacing the inconsistent caption-based alternatives for the first time in image-to-image retrieval evaluation. Finally, we validate the generalizability of our method by applying it to unannotated datasets via automated scene graph generation, while substantially contributing in advancing state-of-the-art in counterfactual image retrieval.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15867
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval
Chaidos, Nikolaos
Dimitriou, Angeliki
Lymperaiou, Maria
Stamou, Giorgos
Computer Vision and Pattern Recognition
Machine Learning
Despite the dominance of convolutional and transformer-based architectures in image-to-image retrieval, these models are prone to biases arising from low-level visual features, such as color. Recognizing the lack of semantic understanding as a key limitation, we propose a novel scene graph-based retrieval framework that emphasizes semantic content over superficial image characteristics. Prior approaches to scene graph retrieval predominantly rely on supervised Graph Neural Networks (GNNs), which require ground truth graph pairs driven from image captions. However, the inconsistency of caption-based supervision stemming from variable text encodings undermine retrieval reliability. To address these, we present SCENIR, a Graph Autoencoder-based unsupervised retrieval framework, which eliminates the dependence on labeled training data. Our model demonstrates superior performance across metrics and runtime efficiency, outperforming existing vision-based, multimodal, and supervised GNN approaches. We further advocate for Graph Edit Distance (GED) as a deterministic and robust ground truth measure for scene graph similarity, replacing the inconsistent caption-based alternatives for the first time in image-to-image retrieval evaluation. Finally, we validate the generalizability of our method by applying it to unannotated datasets via automated scene graph generation, while substantially contributing in advancing state-of-the-art in counterfactual image retrieval.
title SCENIR: Visual Semantic Clarity through Unsupervised Scene Graph Retrieval
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.15867