Saved in:
Bibliographic Details
Main Authors: Mamtani, Sumit, Thesia, Yash
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.20322
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Fine-grained visual classification aims to recognize objects belonging to many subordinate categories of a supercategory, where appearance alone often fails to distinguish highly similar classes. We propose a unified framework that integrates image, text, and metadata via cross-contrastive pre-training. We first align the three modality encoders in a shared embedding space and then fine-tune the image and metadata encoders for classification. On NABirds, our approach improves over the baseline by 7.83% and achieves 84.44% top-1 accuracy, outperforming strong multimodal methods.