Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bi, Qi, Ji, Wei, Yi, Jingjun, Zhan, Haolan, Xia, Gui-Song
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913908363624448
author Bi, Qi
Ji, Wei
Yi, Jingjun
Zhan, Haolan
Xia, Gui-Song
author_facet Bi, Qi
Ji, Wei
Yi, Jingjun
Zhan, Haolan
Xia, Gui-Song
contents High-quality annotation of fine-grained visual categories demands great expert knowledge, which is taxing and time consuming. Alternatively, learning fine-grained visual representation from enormous unlabeled images (e.g., species, brands) by self-supervised learning becomes a feasible solution. However, recent researches find that existing self-supervised learning methods are less qualified to represent fine-grained categories. The bottleneck lies in that the pre-text representation is built from every patch-wise embedding, while fine-grained categories are only determined by several key patches of an image. In this paper, we propose a Cross-level Multi-instance Distillation (CMD) framework to tackle the challenge. Our key idea is to consider the importance of each image patch in determining the fine-grained pre-text representation by multiple instance learning. To comprehensively learn the relation between informative patches and fine-grained semantics, the multi-instance knowledge distillation is implemented on both the region/image crop pairs from the teacher and student net, and the region-image crops inside the teacher / student net, which we term as intra-level multi-instance distillation and inter-level multi-instance distillation. Extensive experiments on CUB-200-2011, Stanford Cars and FGVC Aircraft show that the proposed method outperforms the contemporary method by upto 10.14% and existing state-of-the-art self-supervised learning approaches by upto 19.78% on both top-1 accuracy and Rank-1 retrieval metric.
format Preprint
id arxiv_https___arxiv_org_abs_2401_08860
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization
Bi, Qi
Ji, Wei
Yi, Jingjun
Zhan, Haolan
Xia, Gui-Song
Computer Vision and Pattern Recognition
High-quality annotation of fine-grained visual categories demands great expert knowledge, which is taxing and time consuming. Alternatively, learning fine-grained visual representation from enormous unlabeled images (e.g., species, brands) by self-supervised learning becomes a feasible solution. However, recent researches find that existing self-supervised learning methods are less qualified to represent fine-grained categories. The bottleneck lies in that the pre-text representation is built from every patch-wise embedding, while fine-grained categories are only determined by several key patches of an image. In this paper, we propose a Cross-level Multi-instance Distillation (CMD) framework to tackle the challenge. Our key idea is to consider the importance of each image patch in determining the fine-grained pre-text representation by multiple instance learning. To comprehensively learn the relation between informative patches and fine-grained semantics, the multi-instance knowledge distillation is implemented on both the region/image crop pairs from the teacher and student net, and the region-image crops inside the teacher / student net, which we term as intra-level multi-instance distillation and inter-level multi-instance distillation. Extensive experiments on CUB-200-2011, Stanford Cars and FGVC Aircraft show that the proposed method outperforms the contemporary method by upto 10.14% and existing state-of-the-art self-supervised learning approaches by upto 19.78% on both top-1 accuracy and Rank-1 retrieval metric.
title Cross-Level Multi-Instance Distillation for Self-Supervised Fine-Grained Visual Categorization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2401.08860