ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866910906821115904 |
|---|---|
| author | Wang, Ziyue Chen, Chi Luo, Fuwen Dong, Yurui Zhang, Yuanchi Xu, Yuzhuang Wang, Xiaolong Li, Peng Liu, Yang |
| author_facet | Wang, Ziyue Chen, Chi Luo, Fuwen Dong, Yurui Zhang, Yuanchi Xu, Yuzhuang Wang, Xiaolong Li, Peng Liu, Yang |
| contents | Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked. To address this gap, we propose a novel benchmark named ActiView to evaluate active perception in MLLMs. We focus on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. Meanwhile, intermediate reasoning behaviors of models are also discussed. Given an image, we restrict the perceptual field of a model, requiring it to actively zoom or shift its perceptual field based on reasoning to answer the question successfully. We conduct extensive evaluation over 30 models, including proprietary and open-source models, and observe that restricted perceptual fields play a significant role in enabling active perception. Results reveal a significant gap in the active perception capability of MLLMs, indicating that this area deserves more attention. We hope that ActiView could help develop methods for MLLMs to understand multimodal inputs in more natural and holistic ways. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04659 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models Wang, Ziyue Chen, Chi Luo, Fuwen Dong, Yurui Zhang, Yuanchi Xu, Yuzhuang Wang, Xiaolong Li, Peng Liu, Yang Computer Vision and Pattern Recognition Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked. To address this gap, we propose a novel benchmark named ActiView to evaluate active perception in MLLMs. We focus on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. Meanwhile, intermediate reasoning behaviors of models are also discussed. Given an image, we restrict the perceptual field of a model, requiring it to actively zoom or shift its perceptual field based on reasoning to answer the question successfully. We conduct extensive evaluation over 30 models, including proprietary and open-source models, and observe that restricted perceptual fields play a significant role in enabling active perception. Results reveal a significant gap in the active perception capability of MLLMs, indicating that this area deserves more attention. We hope that ActiView could help develop methods for MLLMs to understand multimodal inputs in more natural and holistic ways. |
| title | ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2410.04659 |