3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Wufei, Chen, Haoyu, Zhang, Guofeng, Chou, Yu-Cheng, Chen, Jieneng, de Melo, Celso M, Yuille, Alan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915495811219456
author Ma, Wufei
Chen, Haoyu
Zhang, Guofeng
Chou, Yu-Cheng
Chen, Jieneng
de Melo, Celso M
Yuille, Alan
author_facet Ma, Wufei
Chen, Haoyu
Zhang, Guofeng
Chou, Yu-Cheng
Chen, Jieneng
de Melo, Celso M
Yuille, Alan
contents 3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2412_07825
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
Ma, Wufei
Chen, Haoyu
Zhang, Guofeng
Chou, Yu-Cheng
Chen, Jieneng
de Melo, Celso M
Yuille, Alan
Computer Vision and Pattern Recognition
3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning abilities by balancing data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images from uncommon 6D viewpoints. Our 3DSRBench provide valuable findings and insights about future development of LMMs with strong spatial reasoning abilities. Our project page is available at https://3dsrbench.github.io/.
title 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.07825