How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Shezheng, Li, Xiaopeng, Li, Shasha, Zhao, Shan, Yu, Jie, Ma, Jun, Mao, Xiaoguang, Zhang, Weimin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913639684898816
author Song, Shezheng
Li, Xiaopeng
Li, Shasha
Zhao, Shan
Yu, Jie
Ma, Jun
Mao, Xiaoguang
Zhang, Weimin
author_facet Song, Shezheng
Li, Xiaopeng
Li, Shasha
Zhao, Shan
Yu, Jie
Ma, Jun
Mao, Xiaoguang
Zhang, Weimin
contents We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, MLLMs still face challenges in addressing the semantic gap in multimodal data, which may lead to erroneous outputs, posing potential risks to society. Selecting the appropriate modality alignment method is crucial, as improper methods might require more parameters without significant performance improvements. This paper aims to explore modality alignment methods for LLMs and their current capabilities. Implementing effective modality alignment can help LLMs address environmental issues and enhance accessibility. The study surveys existing modality alignment methods for MLLMs, categorizing them into four groups: (1) Multimodal Converter, which transforms data into a format that LLMs can understand; (2) Multimodal Perceiver, which improves how LLMs percieve different types of data; (3) Tool Learning, which leverages external tools to convert data into a common format, usually text; and (4) Data-Driven Method, which teaches LLMs to understand specific data types within datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2311_07594
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
Song, Shezheng
Li, Xiaopeng
Li, Shasha
Zhao, Shan
Yu, Jie
Ma, Jun
Mao, Xiaoguang
Zhang, Weimin
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, MLLMs still face challenges in addressing the semantic gap in multimodal data, which may lead to erroneous outputs, posing potential risks to society. Selecting the appropriate modality alignment method is crucial, as improper methods might require more parameters without significant performance improvements. This paper aims to explore modality alignment methods for LLMs and their current capabilities. Implementing effective modality alignment can help LLMs address environmental issues and enhance accessibility. The study surveys existing modality alignment methods for MLLMs, categorizing them into four groups: (1) Multimodal Converter, which transforms data into a format that LLMs can understand; (2) Multimodal Perceiver, which improves how LLMs percieve different types of data; (3) Tool Learning, which leverages external tools to convert data into a common format, usually text; and (4) Data-Driven Method, which teaches LLMs to understand specific data types within datasets.
title How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2311.07594