An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shiri, Fatemeh, Guo, Xiao-Yu, Far, Mona Golestan, Yu, Xin, Haffari, Gholamreza, Li, Yuan-Fang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917832738996224
author Shiri, Fatemeh
Guo, Xiao-Yu
Far, Mona Golestan
Yu, Xin
Haffari, Gholamreza
Li, Yuan-Fang
author_facet Shiri, Fatemeh
Guo, Xiao-Yu
Far, Mona Golestan
Yu, Xin
Haffari, Gholamreza
Li, Yuan-Fang
contents Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. % Moreover, spatial reasoning steps are much less accurate than non-spatial ones across MLLMs. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. Spatial-MM benchmark is available at: https://github.com/FatemehShiri/Spatial-MM
format Preprint
id arxiv_https___arxiv_org_abs_2411_06048
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
Shiri, Fatemeh
Guo, Xiao-Yu
Far, Mona Golestan
Yu, Xin
Haffari, Gholamreza
Li, Yuan-Fang
Computer Vision and Pattern Recognition
Artificial Intelligence
Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. % Moreover, spatial reasoning steps are much less accurate than non-spatial ones across MLLMs. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. Spatial-MM benchmark is available at: https://github.com/FatemehShiri/Spatial-MM
title An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.06048