ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Ziyue, Chen, Chi, Luo, Fuwen, Dong, Yurui, Zhang, Yuanchi, Xu, Yuzhuang, Wang, Xiaolong, Li, Peng, Liu, Yang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910906821115904
author Wang, Ziyue
Chen, Chi
Luo, Fuwen
Dong, Yurui
Zhang, Yuanchi
Xu, Yuzhuang
Wang, Xiaolong
Li, Peng
Liu, Yang
author_facet Wang, Ziyue
Chen, Chi
Luo, Fuwen
Dong, Yurui
Zhang, Yuanchi
Xu, Yuzhuang
Wang, Xiaolong
Li, Peng
Liu, Yang
contents Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked. To address this gap, we propose a novel benchmark named ActiView to evaluate active perception in MLLMs. We focus on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. Meanwhile, intermediate reasoning behaviors of models are also discussed. Given an image, we restrict the perceptual field of a model, requiring it to actively zoom or shift its perceptual field based on reasoning to answer the question successfully. We conduct extensive evaluation over 30 models, including proprietary and open-source models, and observe that restricted perceptual fields play a significant role in enabling active perception. Results reveal a significant gap in the active perception capability of MLLMs, indicating that this area deserves more attention. We hope that ActiView could help develop methods for MLLMs to understand multimodal inputs in more natural and holistic ways.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04659
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
Wang, Ziyue
Chen, Chi
Luo, Fuwen
Dong, Yurui
Zhang, Yuanchi
Xu, Yuzhuang
Wang, Xiaolong
Li, Peng
Liu, Yang
Computer Vision and Pattern Recognition
Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked. To address this gap, we propose a novel benchmark named ActiView to evaluate active perception in MLLMs. We focus on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. Meanwhile, intermediate reasoning behaviors of models are also discussed. Given an image, we restrict the perceptual field of a model, requiring it to actively zoom or shift its perceptual field based on reasoning to answer the question successfully. We conduct extensive evaluation over 30 models, including proprietary and open-source models, and observe that restricted perceptual fields play a significant role in enabling active perception. Results reveal a significant gap in the active perception capability of MLLMs, indicating that this area deserves more attention. We hope that ActiView could help develop methods for MLLMs to understand multimodal inputs in more natural and holistic ways.
title ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.04659