A Survey of Multimodal Large Language Model from A Data-centric Perspective

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Tianyi, Liang, Hao, Wan, Binwang, Xu, Yanran, Li, Xi, Li, Shiyu, Yang, Ling, Li, Bozhou, Wang, Yifan, Cui, Bin, Huang, Ping, Shan, Jiulong, He, Conghui, Yuan, Binhang, Zhang, Wentao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917726249811968
author Bai, Tianyi
Liang, Hao
Wan, Binwang
Xu, Yanran
Li, Xi
Li, Shiyu
Yang, Ling
Li, Bozhou
Wang, Yifan
Cui, Bin
Huang, Ping
Shan, Jiulong
He, Conghui
Yuan, Binhang
Zhang, Wentao
author_facet Bai, Tianyi
Liang, Hao
Wan, Binwang
Xu, Yanran
Li, Xi
Li, Shiyu
Yang, Ling
Li, Bozhou
Wang, Yifan
Cui, Bin
Huang, Ping
Shan, Jiulong
He, Conghui
Yuan, Binhang
Zhang, Wentao
contents Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16640
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey of Multimodal Large Language Model from A Data-centric Perspective
Bai, Tianyi
Liang, Hao
Wan, Binwang
Xu, Yanran
Li, Xi
Li, Shiyu
Yang, Ling
Li, Bozhou
Wang, Yifan
Cui, Bin
Huang, Ping
Shan, Jiulong
He, Conghui
Yuan, Binhang
Zhang, Wentao
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.
title A Survey of Multimodal Large Language Model from A Data-centric Perspective
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2405.16640