Imp: Highly Capable Large Multimodal Models for Mobile Devices

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Shao, Zhenwei, Yu, Zhou, Yu, Jun, Ouyang, Xuecheng, Zheng, Lihao, Gai, Zhenbiao, Wang, Mingyang, Ding, Jiajun
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910463821873152
author Shao, Zhenwei
Yu, Zhou
Yu, Jun
Ouyang, Xuecheng
Zheng, Lihao
Gai, Zhenbiao
Wang, Mingyang
Ding, Jiajun
author_facet Shao, Zhenwei
Yu, Zhou
Yu, Jun
Ouyang, Xuecheng
Zheng, Lihao
Gai, Zhenbiao
Wang, Mingyang
Ding, Jiajun
contents By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp -- a family of highly capable LMMs at the 2B-4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12107
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Imp: Highly Capable Large Multimodal Models for Mobile Devices
Shao, Zhenwei
Yu, Zhou
Yu, Jun
Ouyang, Xuecheng
Zheng, Lihao
Gai, Zhenbiao
Wang, Mingyang
Ding, Jiajun
Computer Vision and Pattern Recognition
Computation and Language
By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp -- a family of highly capable LMMs at the 2B-4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s.
title Imp: Highly Capable Large Multimodal Models for Mobile Devices
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2405.12107