Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zheyuan, Hu, Fengyuan, Lee, Jayjun, Shi, Freda, Kordjamshidi, Parisa, Chai, Joyce, Ma, Ziqiao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912332943196160
author Zhang, Zheyuan
Hu, Fengyuan
Lee, Jayjun
Shi, Freda
Kordjamshidi, Parisa
Chai, Joyce
Ma, Ziqiao
author_facet Zhang, Zheyuan
Hu, Fengyuan
Lee, Jayjun
Shi, Freda
Kordjamshidi, Parisa
Chai, Joyce
Ma, Ziqiao
contents Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17385
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
Zhang, Zheyuan
Hu, Fengyuan
Lee, Jayjun
Shi, Freda
Kordjamshidi, Parisa
Chai, Joyce
Ma, Ziqiao
Computation and Language
Computer Vision and Pattern Recognition
Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning.
title Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.17385