Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhong, Wenliang, Wu, Wenyi, Li, Qi, Barton, Rob, Du, Boxin, Sam, Shioulin, Bouyarmane, Karim, Tutar, Ismail, Huang, Junzhou
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910472876326912
author Zhong, Wenliang
Wu, Wenyi
Li, Qi
Barton, Rob
Du, Boxin
Sam, Shioulin
Bouyarmane, Karim
Tutar, Ismail
Huang, Junzhou
author_facet Zhong, Wenliang
Wu, Wenyi
Li, Qi
Barton, Rob
Du, Boxin
Sam, Shioulin
Bouyarmane, Karim
Tutar, Ismail
Huang, Junzhou
contents Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapters. In this paper, we first establish that adapters using query-based Transformers such as Q-former is a simplified Multi-instance Learning method without considering instance heterogeneity/correlation. We then propose a general component termed Multi-instance Visual Prompt Generator (MIVPG) to incorporate enriched visual representations into LLMs by taking advantage of instance correlation between images or patches for the same sample. Quantatitive evaluation on three public vision-language (VL) datasets from different scenarios shows that the proposed MIVPG improves Q-former in main VL tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_02987
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
Zhong, Wenliang
Wu, Wenyi
Li, Qi
Barton, Rob
Du, Boxin
Sam, Shioulin
Bouyarmane, Karim
Tutar, Ismail
Huang, Junzhou
Computer Vision and Pattern Recognition
Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapters. In this paper, we first establish that adapters using query-based Transformers such as Q-former is a simplified Multi-instance Learning method without considering instance heterogeneity/correlation. We then propose a general component termed Multi-instance Visual Prompt Generator (MIVPG) to incorporate enriched visual representations into LLMs by taking advantage of instance correlation between images or patches for the same sample. Quantatitive evaluation on three public vision-language (VL) datasets from different scenarios shows that the proposed MIVPG improves Q-former in main VL tasks.
title Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.02987