FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tong, Bo, Lai, Bokai, Zhou, Yiyi, Luo, Gen, Shen, Yunhang, Li, Ke, Sun, Xiaoshuai, Ji, Rongrong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915049705046016
author Tong, Bo
Lai, Bokai
Zhou, Yiyi
Luo, Gen
Shen, Yunhang
Li, Ke
Sun, Xiaoshuai
Ji, Rongrong
author_facet Tong, Bo
Lai, Bokai
Zhou, Yiyi
Luo, Gen
Shen, Yunhang
Li, Ke
Sun, Xiaoshuai
Ji, Rongrong
contents Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their actual speedup. In this paper, we propose a powerful and fast tiny MLLM called FlashSloth. Different from previous efforts, FlashSloth focuses on improving the descriptive power of visual tokens in the process of compressing their redundant semantics. In particular, FlashSloth introduces embedded visual compression designs to capture both visually salient and instruction-related image information, so as to achieving superior multimodal performance with fewer visual tokens. Extensive experiments are conducted to validate the proposed FlashSloth, and a bunch of tiny but strong MLLMs are also comprehensively compared, e.g., InternVL2, MiniCPM-V2 and Qwen2-VL. The experimental results show that compared with these advanced tiny MLLMs, our FlashSloth can greatly reduce the number of visual tokens, training memory and computation complexity while retaining high performance on various VL tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2412_04317
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
Tong, Bo
Lai, Bokai
Zhou, Yiyi
Luo, Gen
Shen, Yunhang
Li, Ke
Sun, Xiaoshuai
Ji, Rongrong
Computer Vision and Pattern Recognition
Despite a big leap forward in capability, multimodal large language models (MLLMs) tend to behave like a sloth in practical use, i.e., slow response and large latency. Recent efforts are devoted to building tiny MLLMs for better efficiency, but the plethora of visual tokens still used limit their actual speedup. In this paper, we propose a powerful and fast tiny MLLM called FlashSloth. Different from previous efforts, FlashSloth focuses on improving the descriptive power of visual tokens in the process of compressing their redundant semantics. In particular, FlashSloth introduces embedded visual compression designs to capture both visually salient and instruction-related image information, so as to achieving superior multimodal performance with fewer visual tokens. Extensive experiments are conducted to validate the proposed FlashSloth, and a bunch of tiny but strong MLLMs are also comprehensively compared, e.g., InternVL2, MiniCPM-V2 and Qwen2-VL. The experimental results show that compared with these advanced tiny MLLMs, our FlashSloth can greatly reduce the number of visual tokens, training memory and computation complexity while retaining high performance on various VL tasks.
title FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.04317