CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914786509324288 |
|---|---|
| author | Sun, Shuyang Li, Runjia Torr, Philip Gu, Xiuye Li, Siyang |
| author_facet | Sun, Shuyang Li, Runjia Torr, Philip Gu, Xiuye Li, Siyang |
| contents | Existing open-vocabulary image segmentation methods require a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Consequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM's broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counterparts, but also those fine-tuned with millions of data samples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mIoU on Pascal VOC, COCO Object, and Pascal Context. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_07661 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor Sun, Shuyang Li, Runjia Torr, Philip Gu, Xiuye Li, Siyang Computer Vision and Pattern Recognition Computation and Language Machine Learning Multimedia Existing open-vocabulary image segmentation methods require a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Consequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM's broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counterparts, but also those fine-tuned with millions of data samples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mIoU on Pascal VOC, COCO Object, and Pascal Context. |
| title | CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor |
| topic | Computer Vision and Pattern Recognition Computation and Language Machine Learning Multimedia |
| url | https://arxiv.org/abs/2312.07661 |