Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tumu, Akshar, Shinde, Varad, Kordjamshidi, Parisa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909894192398336
author Tumu, Akshar
Shinde, Varad
Kordjamshidi, Parisa
author_facet Tumu, Akshar
Shinde, Varad
Kordjamshidi, Parisa
contents Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering. In this work, we propose using the Referring Expression Comprehension task instead as a platform for the evaluation of spatial reasoning by VLMs. This platform provides the opportunity for a deeper analysis of spatial comprehension and grounding abilities when there is 1) ambiguity in object detection, 2) complex spatial expressions with a longer sentence structure and multiple spatial relations, and 3) expressions with negation ('not'). In our analysis, we use task-specific architectures as well as large VLMs and highlight their strengths and weaknesses in dealing with these specific situations. While all these models face challenges with the task at hand, the relative behaviors depend on the underlying models and the specific categories of spatial semantics (topological, directional, proximal, etc.). Our results highlight these challenges and behaviors and provide insight into research gaps and future directions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_06146
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
Tumu, Akshar
Shinde, Varad
Kordjamshidi, Parisa
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering. In this work, we propose using the Referring Expression Comprehension task instead as a platform for the evaluation of spatial reasoning by VLMs. This platform provides the opportunity for a deeper analysis of spatial comprehension and grounding abilities when there is 1) ambiguity in object detection, 2) complex spatial expressions with a longer sentence structure and multiple spatial relations, and 3) expressions with negation ('not'). In our analysis, we use task-specific architectures as well as large VLMs and highlight their strengths and weaknesses in dealing with these specific situations. While all these models face challenges with the task at hand, the relative behaviors depend on the underlying models and the specific categories of spatial semantics (topological, directional, proximal, etc.). Our results highlight these challenges and behaviors and provide insight into research gaps and future directions.
title Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.06146