Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Eldesokey, Abdelrahman, Cvejic, Aleksandar, Ghanem, Bernard, Wonka, Peter
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908560673210368
author Eldesokey, Abdelrahman
Cvejic, Aleksandar
Ghanem, Bernard
Wonka, Peter
author_facet Eldesokey, Abdelrahman
Cvejic, Aleksandar
Ghanem, Bernard
Wonka, Peter
contents We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While diffusion model backbones are known to encode semantically rich features, they must also contain visual features to support their image synthesis capabilities. However, isolating these visual features is challenging due to the absence of annotated datasets. To address this, we introduce an automated pipeline that constructs image pairs with annotated semantic and visual correspondences based on existing subject-driven image generation datasets, and design a contrastive architecture to separate the two feature types. Leveraging the disentangled representations, we propose a new metric, Visual Semantic Matching (VSM), that quantifies visual inconsistencies in subject-driven image generation. Empirical results show that our approach outperforms global feature-based metrics such as CLIP, DINO, and vision--language models in quantifying visual inconsistencies while also enabling spatial localization of inconsistent regions. To our knowledge, this is the first method that supports both quantification and localization of inconsistencies in subject-driven generation, offering a valuable tool for advancing this task. Project Page:https://abdo-eldesokey.github.io/mind-the-glitch/
format Preprint
id arxiv_https___arxiv_org_abs_2509_21989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
Eldesokey, Abdelrahman
Cvejic, Aleksandar
Ghanem, Bernard
Wonka, Peter
Computer Vision and Pattern Recognition
We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While diffusion model backbones are known to encode semantically rich features, they must also contain visual features to support their image synthesis capabilities. However, isolating these visual features is challenging due to the absence of annotated datasets. To address this, we introduce an automated pipeline that constructs image pairs with annotated semantic and visual correspondences based on existing subject-driven image generation datasets, and design a contrastive architecture to separate the two feature types. Leveraging the disentangled representations, we propose a new metric, Visual Semantic Matching (VSM), that quantifies visual inconsistencies in subject-driven image generation. Empirical results show that our approach outperforms global feature-based metrics such as CLIP, DINO, and vision--language models in quantifying visual inconsistencies while also enabling spatial localization of inconsistent regions. To our knowledge, this is the first method that supports both quantification and localization of inconsistencies in subject-driven generation, offering a valuable tool for advancing this task. Project Page:https://abdo-eldesokey.github.io/mind-the-glitch/
title Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.21989