CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wysoczańska, Monika, Siméoni, Oriane, Ramamonjisoa, Michaël, Bursuc, Andrei, Trzciński, Tomasz, Pérez, Patrick
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911815468843008
author Wysoczańska, Monika
Siméoni, Oriane
Ramamonjisoa, Michaël
Bursuc, Andrei
Trzciński, Tomasz
Pérez, Patrick
author_facet Wysoczańska, Monika
Siméoni, Oriane
Ramamonjisoa, Michaël
Bursuc, Andrei
Trzciński, Tomasz
Pérez, Patrick
contents The popular CLIP model displays impressive zero-shot capabilities thanks to its seamless interaction with arbitrary text prompts. However, its lack of spatial awareness makes it unsuitable for dense computer vision tasks, e.g., semantic segmentation, without an additional fine-tuning step that often uses annotations and can potentially suppress its original open-vocabulary properties. Meanwhile, self-supervised representation methods have demonstrated good localization properties without human-made annotations nor explicit supervision. In this work, we take the best of both worlds and propose an open-vocabulary semantic segmentation method, which does not require any annotations. We propose to locally improve dense MaskCLIP features, which are computed with a simple modification of CLIP's last pooling layer, by integrating localization priors extracted from self-supervised features. By doing so, we greatly improve the performance of MaskCLIP and produce smooth outputs. Moreover, we show that the used self-supervised feature properties can directly be learnt from CLIP features. Our method CLIP-DINOiser needs only a single forward pass of CLIP and two light convolutional layers at inference, no extra supervision nor extra memory and reaches state-of-the-art results on challenging and fine-grained benchmarks such as COCO, Pascal Context, Cityscapes and ADE20k. The code to reproduce our results is available at https://github.com/wysoczanska/clip_dinoiser.
format Preprint
id arxiv_https___arxiv_org_abs_2312_12359
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
Wysoczańska, Monika
Siméoni, Oriane
Ramamonjisoa, Michaël
Bursuc, Andrei
Trzciński, Tomasz
Pérez, Patrick
Computer Vision and Pattern Recognition
The popular CLIP model displays impressive zero-shot capabilities thanks to its seamless interaction with arbitrary text prompts. However, its lack of spatial awareness makes it unsuitable for dense computer vision tasks, e.g., semantic segmentation, without an additional fine-tuning step that often uses annotations and can potentially suppress its original open-vocabulary properties. Meanwhile, self-supervised representation methods have demonstrated good localization properties without human-made annotations nor explicit supervision. In this work, we take the best of both worlds and propose an open-vocabulary semantic segmentation method, which does not require any annotations. We propose to locally improve dense MaskCLIP features, which are computed with a simple modification of CLIP's last pooling layer, by integrating localization priors extracted from self-supervised features. By doing so, we greatly improve the performance of MaskCLIP and produce smooth outputs. Moreover, we show that the used self-supervised feature properties can directly be learnt from CLIP features. Our method CLIP-DINOiser needs only a single forward pass of CLIP and two light convolutional layers at inference, no extra supervision nor extra memory and reaches state-of-the-art results on challenging and fine-grained benchmarks such as COCO, Pascal Context, Cityscapes and ADE20k. The code to reproduce our results is available at https://github.com/wysoczanska/clip_dinoiser.
title CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.12359