MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916116992884736 |
|---|---|
| author | Chu, Xiangxiang Qiao, Limeng Zhang, Xinyu Xu, Shuang Wei, Fei Yang, Yang Sun, Xiaofei Hu, Yiming Lin, Xinyang Zhang, Bo Shen, Chunhua |
| author_facet | Chu, Xiangxiang Qiao, Limeng Zhang, Xinyu Xu, Shuang Wei, Fei Yang, Yang Sun, Xiaofei Hu, Yiming Lin, Xinyang Zhang, Bo Shen, Chunhua |
| contents | We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM . |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_03766 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Chu, Xiangxiang Qiao, Limeng Zhang, Xinyu Xu, Shuang Wei, Fei Yang, Yang Sun, Xiaofei Hu, Yiming Lin, Xinyang Zhang, Bo Shen, Chunhua Computer Vision and Pattern Recognition Artificial Intelligence We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM . |
| title | MobileVLM V2: Faster and Stronger Baseline for Vision Language Model |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2402.03766 |