Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benschop, Pascal, Dauwels, Justin, van Gemert, Jan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917217351761920
author Benschop, Pascal
Dauwels, Justin
van Gemert, Jan
author_facet Benschop, Pascal
Dauwels, Justin
van Gemert, Jan
contents Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness (recognizing whether an interaction is harmful or benign) and spatial awareness (tracking who does what to whom, and reasoning about relative positions and motion). Through minimal video pairs, we test three challenges: distinguishing violence from benign activity, binding assailant roles across viewpoints, and judging fine-grained trajectory alignment. While we evaluate recent VLMs in a training-free setting, the benchmark is applicable to any video classification model. Results show performance only slightly above chance across tasks. A simple aid, stable color cues, partly reduces assailant role confusions but does not resolve the underlying weakness. By releasing data and code, we aim to provide reproducible diagnostics and seed exploration of lightweight spatial priors to complement large-scale pretraining.
format Preprint
id arxiv_https___arxiv_org_abs_2601_15780
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
Benschop, Pascal
Dauwels, Justin
van Gemert, Jan
Computer Vision and Pattern Recognition
Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness (recognizing whether an interaction is harmful or benign) and spatial awareness (tracking who does what to whom, and reasoning about relative positions and motion). Through minimal video pairs, we test three challenges: distinguishing violence from benign activity, binding assailant roles across viewpoints, and judging fine-grained trajectory alignment. While we evaluate recent VLMs in a training-free setting, the benchmark is applicable to any video classification model. Results show performance only slightly above chance across tasks. A simple aid, stable color cues, partly reduces assailant role confusions but does not resolve the underlying weakness. By releasing data and code, we aim to provide reproducible diagnostics and seed exploration of lightweight spatial priors to complement large-scale pretraining.
title Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.15780