Understanding Self-Supervised Pretraining with Part-Aware Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Jie, Qi, Jiyang, Ding, Mingyu, Chen, Xiaokang, Luo, Ping, Wang, Xinggang, Liu, Wenyu, Wang, Leye, Wang, Jingdong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929218663743488
author Zhu, Jie
Qi, Jiyang
Ding, Mingyu
Chen, Xiaokang
Luo, Ping
Wang, Xinggang
Liu, Wenyu
Wang, Leye
Wang, Jingdong
author_facet Zhu, Jie
Qi, Jiyang
Ding, Mingyu
Chen, Xiaokang
Luo, Ping
Wang, Xinggang
Liu, Wenyu
Wang, Leye
Wang, Jingdong
contents In this paper, we are interested in understanding self-supervised pretraining through studying the capability that self-supervised representation pretraining methods learn part-aware representations. The study is mainly motivated by that random views, used in contrastive learning, and random masked (visible) patches, used in masked image modeling, are often about object parts. We explain that contrastive learning is a part-to-whole task: the projection layer hallucinates the whole object representation from the object part representation learned from the encoder, and that masked image modeling is a part-to-part task: the masked patches of the object are hallucinated from the visible patches. The explanation suggests that the self-supervised pretrained encoder is required to understand the object part. We empirically compare the off-the-shelf encoders pretrained with several representative methods on object-level recognition and part-level recognition. The results show that the fully-supervised model outperforms self-supervised models for object-level recognition, and most self-supervised contrastive learning and masked image modeling methods outperform the fully-supervised method for part-level recognition. It is observed that the combination of contrastive learning and masked image modeling further improves the performance.
format Preprint
id arxiv_https___arxiv_org_abs_2301_11915
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Understanding Self-Supervised Pretraining with Part-Aware Representation Learning
Zhu, Jie
Qi, Jiyang
Ding, Mingyu
Chen, Xiaokang
Luo, Ping
Wang, Xinggang
Liu, Wenyu
Wang, Leye
Wang, Jingdong
Computer Vision and Pattern Recognition
In this paper, we are interested in understanding self-supervised pretraining through studying the capability that self-supervised representation pretraining methods learn part-aware representations. The study is mainly motivated by that random views, used in contrastive learning, and random masked (visible) patches, used in masked image modeling, are often about object parts. We explain that contrastive learning is a part-to-whole task: the projection layer hallucinates the whole object representation from the object part representation learned from the encoder, and that masked image modeling is a part-to-part task: the masked patches of the object are hallucinated from the visible patches. The explanation suggests that the self-supervised pretrained encoder is required to understand the object part. We empirically compare the off-the-shelf encoders pretrained with several representative methods on object-level recognition and part-level recognition. The results show that the fully-supervised model outperforms self-supervised models for object-level recognition, and most self-supervised contrastive learning and masked image modeling methods outperform the fully-supervised method for part-level recognition. It is observed that the combination of contrastive learning and masked image modeling further improves the performance.
title Understanding Self-Supervised Pretraining with Part-Aware Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2301.11915