LocCa: Visual Pretraining with Location-aware Captioners
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916477347561472 |
|---|---|
| author | Wan, Bo Tschannen, Michael Xian, Yongqin Pavetic, Filip Alabdulmohsin, Ibrahim Wang, Xiao Pinto, André Susano Steiner, Andreas Beyer, Lucas Zhai, Xiaohua |
| author_facet | Wan, Bo Tschannen, Michael Xian, Yongqin Pavetic, Filip Alabdulmohsin, Ibrahim Wang, Xiao Pinto, André Susano Steiner, Andreas Beyer, Lucas Zhai, Xiaohua |
| contents | Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretraining remains an area with limited research. In this paper, we propose a simple visual pretraining method with location-aware captioners (LocCa). LocCa uses a simple image captioner task interface, to teach a model to read out rich information, i.e. bounding box coordinates, and captions, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can easily handle multiple tasks during pretraining. Our experiments demonstrate that LocCa outperforms standard captioners significantly on localization downstream tasks while maintaining comparable performance on holistic tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2403_19596 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | LocCa: Visual Pretraining with Location-aware Captioners Wan, Bo Tschannen, Michael Xian, Yongqin Pavetic, Filip Alabdulmohsin, Ibrahim Wang, Xiao Pinto, André Susano Steiner, Andreas Beyer, Lucas Zhai, Xiaohua Computer Vision and Pattern Recognition Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretraining remains an area with limited research. In this paper, we propose a simple visual pretraining method with location-aware captioners (LocCa). LocCa uses a simple image captioner task interface, to teach a model to read out rich information, i.e. bounding box coordinates, and captions, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can easily handle multiple tasks during pretraining. Our experiments demonstrate that LocCa outperforms standard captioners significantly on localization downstream tasks while maintaining comparable performance on holistic tasks. |
| title | LocCa: Visual Pretraining with Location-aware Captioners |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2403.19596 |