HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tao, Chenxin, Su, Shiqian, Zhu, Xizhou, Zhang, Chenyu, Chen, Zhe, Liu, Jiawen, Wang, Wenhai, Lu, Lewei, Huang, Gao, Qiao, Yu, Dai, Jifeng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917917258416128
author Tao, Chenxin
Su, Shiqian
Zhu, Xizhou
Zhang, Chenyu
Chen, Zhe
Liu, Jiawen
Wang, Wenhai
Lu, Lewei
Huang, Gao
Qiao, Yu
Dai, Jifeng
author_facet Tao, Chenxin
Su, Shiqian
Zhu, Xizhou
Zhang, Chenyu
Chen, Zhe
Liu, Jiawen
Wang, Wenhai
Lu, Lewei
Huang, Gao
Qiao, Yu
Dai, Jifeng
contents The rapid advance of Large Language Models (LLMs) has catalyzed the development of Vision-Language Models (VLMs). Monolithic VLMs, which avoid modality-specific encoders, offer a promising alternative to the compositional ones but face the challenge of inferior performance. Most existing monolithic VLMs require tuning pre-trained LLMs to acquire vision abilities, which may degrade their language capabilities. To address this dilemma, this paper presents a novel high-performance monolithic VLM named HoVLE. We note that LLMs have been shown capable of interpreting images, when image embeddings are aligned with text embeddings. The challenge for current monolithic VLMs actually lies in the lack of a holistic embedding module for both vision and language inputs. Therefore, HoVLE introduces a holistic embedding module that converts visual and textual inputs into a shared space, allowing LLMs to process images in the same way as texts. Furthermore, a multi-stage training strategy is carefully designed to empower the holistic embedding module. It is first trained to distill visual features from a pre-trained vision encoder and text embeddings from the LLM, enabling large-scale training with unpaired random images and text tokens. The whole model further undergoes next-token prediction on multi-modal data to align the embeddings. Finally, an instruction-tuning stage is incorporated. Our experiments show that HoVLE achieves performance close to leading compositional models on various benchmarks, outperforming previous monolithic models by a large margin. Model available at https://huggingface.co/OpenGVLab/HoVLE.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16158
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
Tao, Chenxin
Su, Shiqian
Zhu, Xizhou
Zhang, Chenyu
Chen, Zhe
Liu, Jiawen
Wang, Wenhai
Lu, Lewei
Huang, Gao
Qiao, Yu
Dai, Jifeng
Computer Vision and Pattern Recognition
The rapid advance of Large Language Models (LLMs) has catalyzed the development of Vision-Language Models (VLMs). Monolithic VLMs, which avoid modality-specific encoders, offer a promising alternative to the compositional ones but face the challenge of inferior performance. Most existing monolithic VLMs require tuning pre-trained LLMs to acquire vision abilities, which may degrade their language capabilities. To address this dilemma, this paper presents a novel high-performance monolithic VLM named HoVLE. We note that LLMs have been shown capable of interpreting images, when image embeddings are aligned with text embeddings. The challenge for current monolithic VLMs actually lies in the lack of a holistic embedding module for both vision and language inputs. Therefore, HoVLE introduces a holistic embedding module that converts visual and textual inputs into a shared space, allowing LLMs to process images in the same way as texts. Furthermore, a multi-stage training strategy is carefully designed to empower the holistic embedding module. It is first trained to distill visual features from a pre-trained vision encoder and text embeddings from the LLM, enabling large-scale training with unpaired random images and text tokens. The whole model further undergoes next-token prediction on multi-modal data to align the embeddings. Finally, an instruction-tuning stage is incorporated. Our experiments show that HoVLE achieves performance close to leading compositional models on various benchmarks, outperforming previous monolithic models by a large margin. Model available at https://huggingface.co/OpenGVLab/HoVLE.
title HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.16158