Enhancing Vision-Language Pre-training with Rich Supervisions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Yuan, Shi, Kunyu, Zhu, Pengkai, Belval, Edouard, Nuriel, Oren, Appalaraju, Srikar, Ghadar, Shabnam, Mahadevan, Vijay, Tu, Zhuowen, Soatto, Stefano
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915195081719808
author Gao, Yuan
Shi, Kunyu
Zhu, Pengkai
Belval, Edouard
Nuriel, Oren
Appalaraju, Srikar
Ghadar, Shabnam
Mahadevan, Vijay
Tu, Zhuowen
Soatto, Stefano
author_facet Gao, Yuan
Shi, Kunyu
Zhu, Pengkai
Belval, Edouard
Nuriel, Oren
Appalaraju, Srikar
Ghadar, Shabnam
Mahadevan, Vijay
Tu, Zhuowen
Soatto, Stefano
contents We propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble downstream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1% on Widget Captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03346
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Vision-Language Pre-training with Rich Supervisions
Gao, Yuan
Shi, Kunyu
Zhu, Pengkai
Belval, Edouard
Nuriel, Oren
Appalaraju, Srikar
Ghadar, Shabnam
Mahadevan, Vijay
Tu, Zhuowen
Soatto, Stefano
Computer Vision and Pattern Recognition
We propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble downstream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1% on Widget Captioning.
title Enhancing Vision-Language Pre-training with Rich Supervisions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.03346