LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Yichen, Zhu, Minjie, Liu, Ning, Ou, Zhicai, Mou, Xiaofeng, Tang, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909116439461888
author Zhu, Yichen
Zhu, Minjie
Liu, Ning
Ou, Zhicai
Mou, Xiaofeng
Tang, Jian
author_facet Zhu, Yichen
Zhu, Minjie
Liu, Ning
Ou, Zhicai
Mou, Xiaofeng
Tang, Jian
contents In this paper, we introduce LLaVA-$ϕ$ (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate multi-modal dialogues. LLaVA-Phi marks a notable advancement in the realm of compact multi-modal models. It demonstrates that even smaller language models, with as few as 2.7B parameters, can effectively engage in intricate dialogues that integrate both textual and visual elements, provided they are trained with high-quality corpora. Our model delivers commendable performance on publicly available benchmarks that encompass visual comprehension, reasoning, and knowledge-based perception. Beyond its remarkable performance in multi-modal dialogue tasks, our model opens new avenues for applications in time-sensitive environments and systems that require real-time interaction, such as embodied agents. It highlights the potential of smaller language models to achieve sophisticated levels of understanding and interaction, while maintaining greater resource efficiency.The project is available at {https://github.com/zhuyiche/llava-phi}.
format Preprint
id arxiv_https___arxiv_org_abs_2401_02330
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Zhu, Yichen
Zhu, Minjie
Liu, Ning
Ou, Zhicai
Mou, Xiaofeng
Tang, Jian
Computer Vision and Pattern Recognition
Computation and Language
In this paper, we introduce LLaVA-$ϕ$ (LLaVA-Phi), an efficient multi-modal assistant that harnesses the power of the recently advanced small language model, Phi-2, to facilitate multi-modal dialogues. LLaVA-Phi marks a notable advancement in the realm of compact multi-modal models. It demonstrates that even smaller language models, with as few as 2.7B parameters, can effectively engage in intricate dialogues that integrate both textual and visual elements, provided they are trained with high-quality corpora. Our model delivers commendable performance on publicly available benchmarks that encompass visual comprehension, reasoning, and knowledge-based perception. Beyond its remarkable performance in multi-modal dialogue tasks, our model opens new avenues for applications in time-sensitive environments and systems that require real-time interaction, such as embodied agents. It highlights the potential of smaller language models to achieve sophisticated levels of understanding and interaction, while maintaining greater resource efficiency.The project is available at {https://github.com/zhuyiche/llava-phi}.
title LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2401.02330