LocCa: Visual Pretraining with Location-aware Captioners

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wan, Bo, Tschannen, Michael, Xian, Yongqin, Pavetic, Filip, Alabdulmohsin, Ibrahim, Wang, Xiao, Pinto, André Susano, Steiner, Andreas, Beyer, Lucas, Zhai, Xiaohua
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916477347561472
author Wan, Bo
Tschannen, Michael
Xian, Yongqin
Pavetic, Filip
Alabdulmohsin, Ibrahim
Wang, Xiao
Pinto, André Susano
Steiner, Andreas
Beyer, Lucas
Zhai, Xiaohua
author_facet Wan, Bo
Tschannen, Michael
Xian, Yongqin
Pavetic, Filip
Alabdulmohsin, Ibrahim
Wang, Xiao
Pinto, André Susano
Steiner, Andreas
Beyer, Lucas
Zhai, Xiaohua
contents Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretraining remains an area with limited research. In this paper, we propose a simple visual pretraining method with location-aware captioners (LocCa). LocCa uses a simple image captioner task interface, to teach a model to read out rich information, i.e. bounding box coordinates, and captions, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can easily handle multiple tasks during pretraining. Our experiments demonstrate that LocCa outperforms standard captioners significantly on localization downstream tasks while maintaining comparable performance on holistic tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LocCa: Visual Pretraining with Location-aware Captioners
Wan, Bo
Tschannen, Michael
Xian, Yongqin
Pavetic, Filip
Alabdulmohsin, Ibrahim
Wang, Xiao
Pinto, André Susano
Steiner, Andreas
Beyer, Lucas
Zhai, Xiaohua
Computer Vision and Pattern Recognition
Image captioning has been shown as an effective pretraining method similar to contrastive pretraining. However, the incorporation of location-aware information into visual pretraining remains an area with limited research. In this paper, we propose a simple visual pretraining method with location-aware captioners (LocCa). LocCa uses a simple image captioner task interface, to teach a model to read out rich information, i.e. bounding box coordinates, and captions, conditioned on the image pixel input. Thanks to the multitask capabilities of an encoder-decoder architecture, we show that an image captioner can easily handle multiple tasks during pretraining. Our experiments demonstrate that LocCa outperforms standard captioners significantly on localization downstream tasks while maintaining comparable performance on holistic tasks.
title LocCa: Visual Pretraining with Location-aware Captioners
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.19596