Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singh, Simranjit, Pavlakos, Georgios, Stamoulis, Dimitrios
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917677964984320
author Singh, Simranjit
Pavlakos, Georgios
Stamoulis, Dimitrios
author_facet Singh, Simranjit
Pavlakos, Georgios
Stamoulis, Dimitrios
contents As interest in "reformulating" the 3D Visual Question Answering (VQA) problem in the context of foundation models grows, it is imperative to assess how these new paradigms influence existing closed-vocabulary datasets. In this case study, we evaluate the zero-shot performance of foundational models (GPT-4 Vision and GPT-4) on well-established 3D VQA benchmarks, namely 3D-VQA and ScanQA. We provide an investigation to contextualize the performance of GPT-based agents relative to traditional modeling approaches. We find that GPT-based agents without any fine-tuning perform on par with the closed vocabulary approaches. Our findings corroborate recent results that "blind" models establish a surprisingly strong baseline in closed-vocabulary settings. We demonstrate that agents benefit significantly from scene-specific vocabulary via in-context textual grounding. By presenting a preliminary comparison with previous baselines, we hope to inform the community's ongoing efforts to refine multi-modal 3D benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18831
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
Singh, Simranjit
Pavlakos, Georgios
Stamoulis, Dimitrios
Computer Vision and Pattern Recognition
Machine Learning
As interest in "reformulating" the 3D Visual Question Answering (VQA) problem in the context of foundation models grows, it is imperative to assess how these new paradigms influence existing closed-vocabulary datasets. In this case study, we evaluate the zero-shot performance of foundational models (GPT-4 Vision and GPT-4) on well-established 3D VQA benchmarks, namely 3D-VQA and ScanQA. We provide an investigation to contextualize the performance of GPT-based agents relative to traditional modeling approaches. We find that GPT-based agents without any fine-tuning perform on par with the closed vocabulary approaches. Our findings corroborate recent results that "blind" models establish a surprisingly strong baseline in closed-vocabulary settings. We demonstrate that agents benefit significantly from scene-specific vocabulary via in-context textual grounding. By presenting a preliminary comparison with previous baselines, we hope to inform the community's ongoing efforts to refine multi-modal 3D benchmarks.
title Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.18831