Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mirjalili, Vahid, Giahi, Ramin, Kollipara, Sriram, Kekuda, Akshay, Yao, Kehui, Zhao, Kai, Xu, Jianpeng, Nag, Kaushiki, Subramaniam, Sinduja, Biswas, Topojoy, Korpeoglu, Evren, Achan, Kannan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918148713742336
author Mirjalili, Vahid
Giahi, Ramin
Kollipara, Sriram
Kekuda, Akshay
Yao, Kehui
Zhao, Kai
Xu, Jianpeng
Nag, Kaushiki
Subramaniam, Sinduja
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
author_facet Mirjalili, Vahid
Giahi, Ramin
Kollipara, Sriram
Kekuda, Akshay
Yao, Kehui
Zhao, Kai
Xu, Jianpeng
Nag, Kaushiki
Subramaniam, Sinduja
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
contents Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize localization accuracy rather than whether models capture how objects are arranged and related within a scene. This gap is consequential; effective scene understanding requires not only identifying objects, but reasoning about their relative positions, groupings, and depth. In this paper, we present a systematic benchmark for object-centric spatial reasoning in foundation models. Using a controlled synthetic dataset, we evaluate state-of-the-art vision models (e.g., GroundingDINO, Florence-2, OWLv2) and large VLMs (e.g., InternVL, LLaVA, GPT-4o) across three tasks: spatial localization, spatial reasoning, and downstream retrieval tasks. We find a stable trade-off: detectors such as GroundingDINO and OWLv2 deliver precise boxes with limited relational reasoning, while VLMs like SmolVLM and GPT-4o provide coarse layout cues and fluent captions but struggle with fine-grained spatial context. Our study highlights the gap between localization and true spatial understanding, and pointing toward the need for spatially-aware foundation models in the community.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21922
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
Mirjalili, Vahid
Giahi, Ramin
Kollipara, Sriram
Kekuda, Akshay
Yao, Kehui
Zhao, Kai
Xu, Jianpeng
Nag, Kaushiki
Subramaniam, Sinduja
Biswas, Topojoy
Korpeoglu, Evren
Achan, Kannan
Computer Vision and Pattern Recognition
Spatial understanding is a critical capability for vision foundation models. While recent advances in large vision models or vision-language models (VLMs) have expanded recognition capabilities, most benchmarks emphasize localization accuracy rather than whether models capture how objects are arranged and related within a scene. This gap is consequential; effective scene understanding requires not only identifying objects, but reasoning about their relative positions, groupings, and depth. In this paper, we present a systematic benchmark for object-centric spatial reasoning in foundation models. Using a controlled synthetic dataset, we evaluate state-of-the-art vision models (e.g., GroundingDINO, Florence-2, OWLv2) and large VLMs (e.g., InternVL, LLaVA, GPT-4o) across three tasks: spatial localization, spatial reasoning, and downstream retrieval tasks. We find a stable trade-off: detectors such as GroundingDINO and OWLv2 deliver precise boxes with limited relational reasoning, while VLMs like SmolVLM and GPT-4o provide coarse layout cues and fluent captions but struggle with fine-grained spatial context. Our study highlights the gap between localization and true spatial understanding, and pointing toward the need for spatially-aware foundation models in the community.
title Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.21922