Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jian, Yiren, Liu, Tingkai, Tao, Yunzhe, Zhang, Chunhui, Vosoughi, Soroush, Yang, Hongxia
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913238255403008
author Jian, Yiren
Liu, Tingkai
Tao, Yunzhe
Zhang, Chunhui
Vosoughi, Soroush
Yang, Hongxia
author_facet Jian, Yiren
Liu, Tingkai
Tao, Yunzhe
Zhang, Chunhui
Vosoughi, Soroush
Yang, Hongxia
contents In this paper, we introduce $\text{EVL}_{\text{Gen}}$, a streamlined framework designed for the pre-training of visually conditioned language generation models with high computational demands, utilizing frozen pre-trained large language models (LLMs). The conventional approach in vision-language pre-training (VLP) typically involves a two-stage optimization process: an initial resource-intensive phase dedicated to general-purpose vision-language representation learning, focused on extracting and consolidating relevant visual features. This is followed by a subsequent phase that emphasizes end-to-end alignment between visual and linguistic modalities. Our novel one-stage, single-loss framework bypasses the computationally demanding first training stage by gradually merging similar visual tokens during training, while avoiding model collapse caused by single-stage training of BLIP-2 type models. The gradual merging process effectively condenses visual information while preserving semantic richness, resulting in rapid convergence without compromising performance. Our experimental findings demonstrate that our approach accelerates the training of vision-language models by a factor of 5 without a noticeable impact on overall performance. Furthermore, we illustrate that our models significantly narrow the performance gap to current vision-language models using only 1/10 of the data. Finally, we showcase how our image-text models can seamlessly adapt to video-conditioned language generation tasks through novel soft attentive temporal token contextualizing modules. Code is available at \url{https://github.com/yiren-jian/EVLGen}.
format Preprint
id arxiv_https___arxiv_org_abs_2310_03291
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
Jian, Yiren
Liu, Tingkai
Tao, Yunzhe
Zhang, Chunhui
Vosoughi, Soroush
Yang, Hongxia
Computer Vision and Pattern Recognition
In this paper, we introduce $\text{EVL}_{\text{Gen}}$, a streamlined framework designed for the pre-training of visually conditioned language generation models with high computational demands, utilizing frozen pre-trained large language models (LLMs). The conventional approach in vision-language pre-training (VLP) typically involves a two-stage optimization process: an initial resource-intensive phase dedicated to general-purpose vision-language representation learning, focused on extracting and consolidating relevant visual features. This is followed by a subsequent phase that emphasizes end-to-end alignment between visual and linguistic modalities. Our novel one-stage, single-loss framework bypasses the computationally demanding first training stage by gradually merging similar visual tokens during training, while avoiding model collapse caused by single-stage training of BLIP-2 type models. The gradual merging process effectively condenses visual information while preserving semantic richness, resulting in rapid convergence without compromising performance. Our experimental findings demonstrate that our approach accelerates the training of vision-language models by a factor of 5 without a noticeable impact on overall performance. Furthermore, we illustrate that our models significantly narrow the performance gap to current vision-language models using only 1/10 of the data. Finally, we showcase how our image-text models can seamlessly adapt to video-conditioned language generation tasks through novel soft attentive temporal token contextualizing modules. Code is available at \url{https://github.com/yiren-jian/EVLGen}.
title Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.03291