OpenView: Empowering MLLMs with Out-of-view VQA

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Qixiang, Zhang, Cheng, Fu, Chi-Wing, Ye, Jingwen, Cai, Jianfei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911329575501824
author Chen, Qixiang
Zhang, Cheng
Fu, Chi-Wing
Ye, Jingwen
Cai, Jianfei
author_facet Chen, Qixiang
Zhang, Cheng
Fu, Chi-Wing
Ye, Jingwen
Cai, Jianfei
contents Recent multimodal large language models (MLLMs) show great potential in natural image understanding. Yet, they perform well, mainly on reasoning in-view contents within the image frame. This paper presents the first study on out-of-view (OOV) understanding, i.e., the ability to reason objects, activities, and scenes beyond the visible frame of a perspective view. Our technical contributions are threefold. First, we design OpenView, a four-stage pipeline to massively generate multi-choice VQA by leveraging panoramic imagery to enable context-rich and spatial-grounded VQA synthesis with free-view framing. Second, we curate OpenView-Dataset, a high-quality synthetic dataset from diverse real-world panoramas to empower MLLMs upon supervised fine-tuning. Third, we build OpenView-Bench, a benchmark that jointly measures choice and rationale accuracy for interpretable and diagnosable evaluation. Experimental results show that despite having a large gap from human performance in OOV VQA answer selection, upon empowered by OpenView, multiple MLLMs can consistently boost their performance, uplifted from 48.6% to 64.1% on average. Code, benchmark, and data will be available at https://github.com/q1xiangchen/OpenView.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18563
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OpenView: Empowering MLLMs with Out-of-view VQA
Chen, Qixiang
Zhang, Cheng
Fu, Chi-Wing
Ye, Jingwen
Cai, Jianfei
Computer Vision and Pattern Recognition
Recent multimodal large language models (MLLMs) show great potential in natural image understanding. Yet, they perform well, mainly on reasoning in-view contents within the image frame. This paper presents the first study on out-of-view (OOV) understanding, i.e., the ability to reason objects, activities, and scenes beyond the visible frame of a perspective view. Our technical contributions are threefold. First, we design OpenView, a four-stage pipeline to massively generate multi-choice VQA by leveraging panoramic imagery to enable context-rich and spatial-grounded VQA synthesis with free-view framing. Second, we curate OpenView-Dataset, a high-quality synthetic dataset from diverse real-world panoramas to empower MLLMs upon supervised fine-tuning. Third, we build OpenView-Bench, a benchmark that jointly measures choice and rationale accuracy for interpretable and diagnosable evaluation. Experimental results show that despite having a large gap from human performance in OOV VQA answer selection, upon empowered by OpenView, multiple MLLMs can consistently boost their performance, uplifted from 48.6% to 64.1% on average. Code, benchmark, and data will be available at https://github.com/q1xiangchen/OpenView.
title OpenView: Empowering MLLMs with Out-of-view VQA
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.18563