Efficient Multimodal Learning from Data-centric Perspective

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: He, Muyang, Liu, Yexin, Wu, Boya, Yuan, Jianhao, Wang, Yueze, Huang, Tiejun, Zhao, Bo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910536276377600
author He, Muyang
Liu, Yexin
Wu, Boya
Yuan, Jianhao
Wang, Yueze
Huang, Tiejun
Zhao, Bo
author_facet He, Muyang
Liu, Yexin
Wu, Boya
Yuan, Jianhao
Wang, Yueze
Huang, Tiejun
Zhao, Bo
contents Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both training and inference, limiting accessibility to the broader research and user communities. A straightforward solution is to leverage smaller pre-trained vision and language models, which inevitably cause significant performance drops. In this paper, we demonstrate the possibility of training a smaller but better MLLM with high-quality training data. Specifically, we introduce Bunny, a family of lightweight MLLMs with flexible vision and language backbones for efficient multimodal learning from selected training data. Experiments show that our Bunny-4B/8B outperforms the state-of-the-art large MLLMs on multiple benchmarks. We expect that this work can provide the community with a clean and flexible open-source tool for further research and development. The code, models, and data can be found in https://github.com/BAAI-DCAI/Bunny.
format Preprint
id arxiv_https___arxiv_org_abs_2402_11530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Multimodal Learning from Data-centric Perspective
He, Muyang
Liu, Yexin
Wu, Boya
Yuan, Jianhao
Wang, Yueze
Huang, Tiejun
Zhao, Bo
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both training and inference, limiting accessibility to the broader research and user communities. A straightforward solution is to leverage smaller pre-trained vision and language models, which inevitably cause significant performance drops. In this paper, we demonstrate the possibility of training a smaller but better MLLM with high-quality training data. Specifically, we introduce Bunny, a family of lightweight MLLMs with flexible vision and language backbones for efficient multimodal learning from selected training data. Experiments show that our Bunny-4B/8B outperforms the state-of-the-art large MLLMs on multiple benchmarks. We expect that this work can provide the community with a clean and flexible open-source tool for further research and development. The code, models, and data can be found in https://github.com/BAAI-DCAI/Bunny.
title Efficient Multimodal Learning from Data-centric Perspective
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.11530