DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Xin, Tang, Hao, Yan, Rui, Tang, Jinhui, Li, Zechao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929326224572416
author Jiang, Xin
Tang, Hao
Yan, Rui
Tang, Jinhui
Li, Zechao
author_facet Jiang, Xin
Tang, Hao
Yan, Rui
Tang, Jinhui
Li, Zechao
contents Fine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings.
format Preprint
id arxiv_https___arxiv_org_abs_2404_15771
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
Jiang, Xin
Tang, Hao
Yan, Rui
Tang, Jinhui
Li, Zechao
Computer Vision and Pattern Recognition
Multimedia
Fine-grained image retrieval (FGIR) is to learn visual representations that distinguish visually similar objects while maintaining generalization. Existing methods propose to generate discriminative features, but rarely consider the particularity of the FGIR task itself. This paper presents a meticulous analysis leading to the proposal of practical guidelines to identify subcategory-specific discrepancies and generate discriminative features to design effective FGIR models. These guidelines include emphasizing the object (G1), highlighting subcategory-specific discrepancies (G2), and employing effective training strategy (G3). Following G1 and G2, we design a novel Dual Visual Filtering mechanism for the plain visual transformer, denoted as DVF, to capture subcategory-specific discrepancies. Specifically, the dual visual filtering mechanism comprises an object-oriented module and a semantic-oriented module. These components serve to magnify objects and identify discriminative regions, respectively. Following G3, we implement a discriminative model training strategy to improve the discriminability and generalization ability of DVF. Extensive analysis and ablation studies confirm the efficacy of our proposed guidelines. Without bells and whistles, the proposed DVF achieves state-of-the-art performance on three widely-used fine-grained datasets in closed-set and open-set settings.
title DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2404.15771