DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Ruohong, Hu, Peng, Li, Yunfan, Peng, Xi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909770593599488
author Yang, Ruohong
Hu, Peng
Li, Yunfan
Peng, Xi
author_facet Yang, Ruohong
Hu, Peng
Li, Yunfan
Peng, Xi
contents Unsupervised cross-domain image retrieval (UCIR) aims to retrieve images of the same category across diverse domains without relying on annotations. Existing UCIR methods, which align cross-domain features for the entire image, often struggle with the domain gap, as the object features critical for retrieval are frequently entangled with domain-specific styles. To address this challenge, we propose DUDE, a novel UCIR method building upon feature disentanglement. In brief, DUDE leverages a text-to-image generative model to disentangle object features from domain-specific styles, thus facilitating semantical image retrieval. To further achieve reliable alignment of the disentangled object features, DUDE aligns mutual neighbors from within domains to across domains in a progressive manner. Extensive experiments demonstrate that DUDE achieves state-of-the-art performance across three benchmark datasets over 13 domains. The code will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04193
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
Yang, Ruohong
Hu, Peng
Li, Yunfan
Peng, Xi
Computer Vision and Pattern Recognition
Machine Learning
Unsupervised cross-domain image retrieval (UCIR) aims to retrieve images of the same category across diverse domains without relying on annotations. Existing UCIR methods, which align cross-domain features for the entire image, often struggle with the domain gap, as the object features critical for retrieval are frequently entangled with domain-specific styles. To address this challenge, we propose DUDE, a novel UCIR method building upon feature disentanglement. In brief, DUDE leverages a text-to-image generative model to disentangle object features from domain-specific styles, thus facilitating semantical image retrieval. To further achieve reliable alignment of the disentangled object features, DUDE aligns mutual neighbors from within domains to across domains in a progressive manner. Extensive experiments demonstrate that DUDE achieves state-of-the-art performance across three benchmark datasets over 13 domains. The code will be released.
title DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.04193