A Survey of Multimodal Large Language Model from A Data-centric Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917726249811968 |
|---|---|
| author | Bai, Tianyi Liang, Hao Wan, Binwang Xu, Yanran Li, Xi Li, Shiyu Yang, Ling Li, Bozhou Wang, Yifan Cui, Bin Huang, Ping Shan, Jiulong He, Conghui Yuan, Binhang Zhang, Wentao |
| author_facet | Bai, Tianyi Liang, Hao Wan, Binwang Xu, Yanran Li, Xi Li, Shiyu Yang, Ling Li, Bozhou Wang, Yifan Cui, Bin Huang, Ping Shan, Jiulong He, Conghui Yuan, Binhang Zhang, Wentao |
| contents | Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_16640 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | A Survey of Multimodal Large Language Model from A Data-centric Perspective Bai, Tianyi Liang, Hao Wan, Binwang Xu, Yanran Li, Xi Li, Shiyu Yang, Ling Li, Bozhou Wang, Yifan Cui, Bin Huang, Ping Shan, Jiulong He, Conghui Yuan, Binhang Zhang, Wentao Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Multimedia Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field. |
| title | A Survey of Multimodal Large Language Model from A Data-centric Perspective |
| topic | Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Multimedia |
| url | https://arxiv.org/abs/2405.16640 |