Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qi, Daiqing, Zhao, Handong, Wei, Zijun, Li, Sheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929588362280960
author Qi, Daiqing
Zhao, Handong
Wei, Zijun
Li, Sheng
author_facet Qi, Daiqing
Zhao, Handong
Wei, Zijun
Li, Sheng
contents Despite recent advances in the general visual instruction-following ability of Multimodal Large Language Models (MLLMs), they still struggle with critical problems when required to provide a precise and detailed response to a visual instruction: (1) failure to identify novel objects or entities, (2) mention of non-existent objects, and (3) neglect of object's attributed details. Intuitive solutions include improving the size and quality of data or using larger foundation models. They show effectiveness in mitigating these issues, but at an expensive cost of collecting a vast amount of new data and introducing a significantly larger model. Standing at the intersection of these approaches, we examine the three object-oriented problems from the perspective of the image-to-text mapping process by the multimodal connector. In this paper, we first identify the limitations of multimodal connectors stemming from insufficient training data. Driven by this, we propose to enhance the mapping with retrieval-augmented tag tokens, which contain rich object-aware information such as object names and attributes. With our Tag-grounded visual instruction tuning with retrieval Augmentation (TUNA), we outperform baselines that share the same language model and training data on 12 benchmarks. Furthermore, we show the zero-shot capability of TUNA when provided with specific datastores.
format Preprint
id arxiv_https___arxiv_org_abs_2406_10839
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
Qi, Daiqing
Zhao, Handong
Wei, Zijun
Li, Sheng
Computer Vision and Pattern Recognition
Computation and Language
Despite recent advances in the general visual instruction-following ability of Multimodal Large Language Models (MLLMs), they still struggle with critical problems when required to provide a precise and detailed response to a visual instruction: (1) failure to identify novel objects or entities, (2) mention of non-existent objects, and (3) neglect of object's attributed details. Intuitive solutions include improving the size and quality of data or using larger foundation models. They show effectiveness in mitigating these issues, but at an expensive cost of collecting a vast amount of new data and introducing a significantly larger model. Standing at the intersection of these approaches, we examine the three object-oriented problems from the perspective of the image-to-text mapping process by the multimodal connector. In this paper, we first identify the limitations of multimodal connectors stemming from insufficient training data. Driven by this, we propose to enhance the mapping with retrieval-augmented tag tokens, which contain rich object-aware information such as object names and attributes. With our Tag-grounded visual instruction tuning with retrieval Augmentation (TUNA), we outperform baselines that share the same language model and training data on 12 benchmarks. Furthermore, we show the zero-shot capability of TUNA when provided with specific datastores.
title Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2406.10839