TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Chaoya, Xu, Haiyang, Li, Chenliang, Yan, Miang, Ye, Wei, Zhang, Shikun, Bi, Bin, Huang, Songfang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909810874646528
author Jiang, Chaoya
Xu, Haiyang
Li, Chenliang
Yan, Miang
Ye, Wei
Zhang, Shikun
Bi, Bin
Huang, Songfang
author_facet Jiang, Chaoya
Xu, Haiyang
Li, Chenliang
Yan, Miang
Ye, Wei
Zhang, Shikun
Bi, Bin
Huang, Songfang
contents Vision Transformers (ViTs) have been widely used in large-scale Vision and Language Pre-training (VLP) models. Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence. To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with \textbf{T}ext-\textbf{R}elevant \textbf{I}mage \textbf{P}atch \textbf{S}election, namely TRIPS, which reduces the visual sequence progressively with a text-guided patch-selection layer in the visual backbone for efficient training and inference. The patch-selection layer can dynamically compute text-dependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner. Meanwhile, TRIPS does not introduce extra parameters to ViTs. Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40\% over previous similar VLP models, yet with competitive or better downstream task performance.
format Preprint
id arxiv_https___arxiv_org_abs_2305_04474
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
Jiang, Chaoya
Xu, Haiyang
Li, Chenliang
Yan, Miang
Ye, Wei
Zhang, Shikun
Bi, Bin
Huang, Songfang
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision Transformers (ViTs) have been widely used in large-scale Vision and Language Pre-training (VLP) models. Though previous VLP works have proved the effectiveness of ViTs, they still suffer from computational efficiency brought by the long visual sequence. To tackle this problem, in this paper, we propose an efficient vision-and-language pre-training model with \textbf{T}ext-\textbf{R}elevant \textbf{I}mage \textbf{P}atch \textbf{S}election, namely TRIPS, which reduces the visual sequence progressively with a text-guided patch-selection layer in the visual backbone for efficient training and inference. The patch-selection layer can dynamically compute text-dependent visual attention to identify the attentive image tokens with text guidance and fuse inattentive ones in an end-to-end manner. Meanwhile, TRIPS does not introduce extra parameters to ViTs. Experimental results on a variety of popular benchmark datasets demonstrate that TRIPS gain a speedup of 40\% over previous similar VLP models, yet with competitive or better downstream task performance.
title TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2305.04474