I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meng, Zaiqiao, Zhou, Hao, Chen, Yifang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914947705864192
author Meng, Zaiqiao
Zhou, Hao
Chen, Yifang
author_facet Meng, Zaiqiao
Zhou, Hao
Chen, Yifang
contents Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. However, existing \VLMs{}' visual spatial reasoning capabilities are often inadequate, struggling even with basic tasks such as distinguishing left from right. To address this, we propose the \ours{} model, designed to enhance the visual spatial reasoning abilities of VLMS. ZeroVLM employs Zero-1-to-3, a 3D reconstruction model for obtaining different views of the input images and incorporates a prompting mechanism to further improve visual spatial reasoning. Experimental results on four visual spatial reasoning datasets show that our \ours{} achieves up to 19.48% accuracy improvement, which indicates the effectiveness of the 3D reconstruction and prompting mechanisms of our ZeroVLM.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14133
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction
Meng, Zaiqiao
Zhou, Hao
Chen, Yifang
Computation and Language
Visual Language Models (VLMs) are essential for various tasks, particularly visual reasoning tasks, due to their robust multi-modal information integration, visual reasoning capabilities, and contextual awareness. However, existing \VLMs{}' visual spatial reasoning capabilities are often inadequate, struggling even with basic tasks such as distinguishing left from right. To address this, we propose the \ours{} model, designed to enhance the visual spatial reasoning abilities of VLMS. ZeroVLM employs Zero-1-to-3, a 3D reconstruction model for obtaining different views of the input images and incorporates a prompting mechanism to further improve visual spatial reasoning. Experimental results on four visual spatial reasoning datasets show that our \ours{} achieves up to 19.48% accuracy improvement, which indicates the effectiveness of the 3D reconstruction and prompting mechanisms of our ZeroVLM.
title I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction
topic Computation and Language
url https://arxiv.org/abs/2407.14133