CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911763692257280 |
|---|---|
| author | Wu, Size Zhang, Wenwei Xu, Lumin Jin, Sheng Li, Xiangtai Liu, Wentao Loy, Chen Change |
| author_facet | Wu, Size Zhang, Wenwei Xu, Lumin Jin, Sheng Li, Xiangtai Liu, Wentao Loy, Chen Change |
| contents | Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP models, particularly those incorporating vision transformers (ViTs), have exhibited remarkable generalization ability in zero-shot image classification. However, when transferring the vision-language alignment of CLIP from global image representation to local region representation for the open-vocabulary dense prediction tasks, CLIP ViTs suffer from the domain shift from full images to local image regions. In this paper, we embark on an in-depth analysis of the region-language alignment in CLIP models, which is essential for downstream open-vocabulary dense prediction tasks. Subsequently, we propose an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs. CLIPSelf empowers ViTs to distill itself by aligning a region representation extracted from its dense feature map with the image-level representation of the corresponding image crop. With the enhanced CLIP ViTs, we achieve new state-of-the-art performance on open-vocabulary object detection, semantic segmentation, and panoptic segmentation across various benchmarks. Models and code are released at https://github.com/wusize/CLIPSelf. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_01403 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction Wu, Size Zhang, Wenwei Xu, Lumin Jin, Sheng Li, Xiangtai Liu, Wentao Loy, Chen Change Computer Vision and Pattern Recognition Open-vocabulary dense prediction tasks including object detection and image segmentation have been advanced by the success of Contrastive Language-Image Pre-training (CLIP). CLIP models, particularly those incorporating vision transformers (ViTs), have exhibited remarkable generalization ability in zero-shot image classification. However, when transferring the vision-language alignment of CLIP from global image representation to local region representation for the open-vocabulary dense prediction tasks, CLIP ViTs suffer from the domain shift from full images to local image regions. In this paper, we embark on an in-depth analysis of the region-language alignment in CLIP models, which is essential for downstream open-vocabulary dense prediction tasks. Subsequently, we propose an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs. CLIPSelf empowers ViTs to distill itself by aligning a region representation extracted from its dense feature map with the image-level representation of the corresponding image crop. With the enhanced CLIP ViTs, we achieve new state-of-the-art performance on open-vocabulary object detection, semantic segmentation, and panoptic segmentation across various benchmarks. Models and code are released at https://github.com/wusize/CLIPSelf. |
| title | CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2310.01403 |