Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lu, Shiyin, Li, Yang, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Ye, Han-Jia
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914836477116416
author Lu, Shiyin
Li, Yang
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Ye, Han-Jia
author_facet Lu, Shiyin
Li, Yang
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Ye, Han-Jia
contents Current Multimodal Large Language Models (MLLMs) typically integrate a pre-trained LLM with another pre-trained vision transformer through a connector, such as an MLP, endowing the LLM with visual capabilities. However, the misalignment between two embedding strategies in MLLMs -- the structural textual embeddings based on an embedding look-up table and the continuous embeddings generated directly by the vision encoder -- makes challenges for a more seamless fusion of visual and textual information. We propose Ovis, a novel MLLM architecture designed to structurally align visual and textual embeddings. Ovis integrates an additional learnable visual embedding table into the visual encoder's process. To capture rich visual semantics, each image patch indexes the visual embedding table multiple times, resulting in a final visual embedding that is a probabilistic combination of the indexed embeddings. This structural approach mirrors the method used for generating textual embeddings. Empirical evaluations on various multimodal benchmarks show that Ovis outperforms open-source MLLMs of similar parameter scales and even surpasses the proprietary model Qwen-VL-Plus overall. These results highlight the potential of Ovis' structured visual representation for advancing MLLM architectural design and promoting more effective multimodal learning. Code, datasets, and models are available at https://github.com/AIDC-AI/Ovis.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20797
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Lu, Shiyin
Li, Yang
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Ye, Han-Jia
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Current Multimodal Large Language Models (MLLMs) typically integrate a pre-trained LLM with another pre-trained vision transformer through a connector, such as an MLP, endowing the LLM with visual capabilities. However, the misalignment between two embedding strategies in MLLMs -- the structural textual embeddings based on an embedding look-up table and the continuous embeddings generated directly by the vision encoder -- makes challenges for a more seamless fusion of visual and textual information. We propose Ovis, a novel MLLM architecture designed to structurally align visual and textual embeddings. Ovis integrates an additional learnable visual embedding table into the visual encoder's process. To capture rich visual semantics, each image patch indexes the visual embedding table multiple times, resulting in a final visual embedding that is a probabilistic combination of the indexed embeddings. This structural approach mirrors the method used for generating textual embeddings. Empirical evaluations on various multimodal benchmarks show that Ovis outperforms open-source MLLMs of similar parameter scales and even surpasses the proprietary model Qwen-VL-Plus overall. These results highlight the potential of Ovis' structured visual representation for advancing MLLM architectural design and promoting more effective multimodal learning. Code, datasets, and models are available at https://github.com/AIDC-AI/Ovis.
title Ovis: Structural Embedding Alignment for Multimodal Large Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.20797