Saved in:
Bibliographic Details
Main Authors: Li, Zhong-Yu, Yin, Bo-Wen, Liu, Yongxiang, Liu, Li, Cheng, Ming-Ming
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2310.05108
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918160346644480
author Li, Zhong-Yu
Yin, Bo-Wen
Liu, Yongxiang
Liu, Li
Cheng, Ming-Ming
author_facet Li, Zhong-Yu
Yin, Bo-Wen
Liu, Yongxiang
Liu, Li
Cheng, Ming-Ming
contents Incorporating heterogeneous representations from different architectures has facilitated various vision tasks, e.g., some hybrid networks combine transformers and convolutions. However, complementarity between such heterogeneous architectures has not been well exploited in self-supervised learning. Thus, we propose Heterogeneous Self-Supervised Learning (HSSL), which enforces a base model to learn from an auxiliary head whose architecture is heterogeneous from the base model. In this process, HSSL endows the base model with new characteristics in a representation learning way without structural changes. To comprehensively understand the HSSL, we conduct experiments on various heterogeneous pairs containing a base model and an auxiliary head. We discover that the representation quality of the base model moves up as their architecture discrepancy grows. This observation motivates us to propose a search strategy that quickly determines the most suitable auxiliary head for a specific base model to learn and several simple but effective methods to enlarge the model discrepancy. The HSSL is compatible with various self-supervised methods, achieving superior performances on various downstream tasks, including image classification, semantic segmentation, instance segmentation, and object detection. The codes are available at https://github.com/NK-JittorCV/Self-Supervised/.
format Preprint
id arxiv_https___arxiv_org_abs_2310_05108
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Enhancing Representations through Heterogeneous Self-Supervised Learning
Li, Zhong-Yu
Yin, Bo-Wen
Liu, Yongxiang
Liu, Li
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
Incorporating heterogeneous representations from different architectures has facilitated various vision tasks, e.g., some hybrid networks combine transformers and convolutions. However, complementarity between such heterogeneous architectures has not been well exploited in self-supervised learning. Thus, we propose Heterogeneous Self-Supervised Learning (HSSL), which enforces a base model to learn from an auxiliary head whose architecture is heterogeneous from the base model. In this process, HSSL endows the base model with new characteristics in a representation learning way without structural changes. To comprehensively understand the HSSL, we conduct experiments on various heterogeneous pairs containing a base model and an auxiliary head. We discover that the representation quality of the base model moves up as their architecture discrepancy grows. This observation motivates us to propose a search strategy that quickly determines the most suitable auxiliary head for a specific base model to learn and several simple but effective methods to enlarge the model discrepancy. The HSSL is compatible with various self-supervised methods, achieving superior performances on various downstream tasks, including image classification, semantic segmentation, instance segmentation, and object detection. The codes are available at https://github.com/NK-JittorCV/Self-Supervised/.
title Enhancing Representations through Heterogeneous Self-Supervised Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.05108