Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jiang, Songtao, Zhang, Yan, Zhou, Chenyi, Jin, Yeying, Feng, Yang, Wu, Jian, Liu, Zuozhu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913302531014656
author Jiang, Songtao
Zhang, Yan
Zhou, Chenyi
Jin, Yeying
Feng, Yang
Wu, Jian
Liu, Zuozhu
author_facet Jiang, Songtao
Zhang, Yan
Zhou, Chenyi
Jin, Yeying
Feng, Yang
Wu, Jian
Liu, Zuozhu
contents Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in object-oriented perception tasks which demand fine-grained understanding of object identities, locations or attributes, as indicated by empirical findings. This is mainly due to their limited capability to effectively integrate complex visual cues with textual information and potential object hallucinations. In this paper, we present a novel approach, Joint Visual and Text Prompting (VTPrompt), that employs fine-grained visual information to enhance the capability of MLLMs in VQA, especially for object-oriented perception. VTPrompt merges visual and text prompts to extract key concepts from textual questions and employs a detection model to highlight relevant objects as visual prompts in images. The processed images alongside text prompts are subsequently fed into MLLMs to produce more accurate answers. Our experiments with GPT-4V and Gemini Pro, on three benchmarks, i.e., MME , MMB and POPE, demonstrate significant improvements. Particularly, our method led to a score improvement of up to 183.5 for GPT-4V on MME and enhanced MMB performance by 8.17\% for GPT-4V and 15.69\% for Gemini Pro.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04514
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
Jiang, Songtao
Zhang, Yan
Zhou, Chenyi
Jin, Yeying
Feng, Yang
Wu, Jian
Liu, Zuozhu
Computation and Language
Multimodal Large Language Models (MLLMs) such as GPT-4V and Gemini Pro face challenges in achieving human-level perception in Visual Question Answering (VQA), particularly in object-oriented perception tasks which demand fine-grained understanding of object identities, locations or attributes, as indicated by empirical findings. This is mainly due to their limited capability to effectively integrate complex visual cues with textual information and potential object hallucinations. In this paper, we present a novel approach, Joint Visual and Text Prompting (VTPrompt), that employs fine-grained visual information to enhance the capability of MLLMs in VQA, especially for object-oriented perception. VTPrompt merges visual and text prompts to extract key concepts from textual questions and employs a detection model to highlight relevant objects as visual prompts in images. The processed images alongside text prompts are subsequently fed into MLLMs to produce more accurate answers. Our experiments with GPT-4V and Gemini Pro, on three benchmarks, i.e., MME , MMB and POPE, demonstrate significant improvements. Particularly, our method led to a score improvement of up to 183.5 for GPT-4V on MME and enhanced MMB performance by 8.17\% for GPT-4V and 15.69\% for Gemini Pro.
title Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2404.04514