MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chu, Xiangxiang, Qiao, Limeng, Zhang, Xinyu, Xu, Shuang, Wei, Fei, Yang, Yang, Sun, Xiaofei, Hu, Yiming, Lin, Xinyang, Zhang, Bo, Shen, Chunhua
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916116992884736
author Chu, Xiangxiang
Qiao, Limeng
Zhang, Xinyu
Xu, Shuang
Wei, Fei
Yang, Yang
Sun, Xiaofei
Hu, Yiming
Lin, Xinyang
Zhang, Bo
Shen, Chunhua
author_facet Chu, Xiangxiang
Qiao, Limeng
Zhang, Xinyu
Xu, Shuang
Wei, Fei
Yang, Yang
Sun, Xiaofei
Hu, Yiming
Lin, Xinyang
Zhang, Bo
Shen, Chunhua
contents We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM .
format Preprint
id arxiv_https___arxiv_org_abs_2402_03766
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Chu, Xiangxiang
Qiao, Limeng
Zhang, Xinyu
Xu, Shuang
Wei, Fei
Yang, Yang
Sun, Xiaofei
Hu, Yiming
Lin, Xinyang
Zhang, Bo
Shen, Chunhua
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM .
title MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2402.03766