VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Shijie, Vilesov, Alexander, He, Xuehai, Wan, Ziyu, Zhang, Shuwang, Nagachandra, Aditya, Chang, Di, Chen, Dongdong, Wang, Xin Eric, Kadambi, Achuta
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912524710969344
author Zhou, Shijie
Vilesov, Alexander
He, Xuehai
Wan, Ziyu
Zhang, Shuwang
Nagachandra, Aditya
Chang, Di
Chen, Dongdong
Wang, Xin Eric
Kadambi, Achuta
author_facet Zhou, Shijie
Vilesov, Alexander
He, Xuehai
Wan, Ziyu
Zhang, Shuwang
Nagachandra, Aditya
Chang, Di
Chen, Dongdong
Wang, Xin Eric
Kadambi, Achuta
contents Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_02095
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
Zhou, Shijie
Vilesov, Alexander
He, Xuehai
Wan, Ziyu
Zhang, Shuwang
Nagachandra, Aditya
Chang, Di
Chen, Dongdong
Wang, Xin Eric
Kadambi, Achuta
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision language models (VLMs) have shown remarkable capabilities in integrating linguistic and visual reasoning but remain fundamentally limited in understanding dynamic spatiotemporal interactions. Humans effortlessly track and reason about object movements, rotations, and perspective shifts-abilities essential for robust dynamic real-world understanding yet notably lacking in current VLMs. In this paper, we introduce VLM4D, the first benchmark specifically designed to evaluate the spatiotemporal reasoning capabilities of VLMs. Our benchmark comprises diverse real-world and synthetic videos accompanied by carefully curated question-answer pairs emphasizing translational and rotational motions, perspective awareness, and motion continuity. Through comprehensive evaluations of state-of-the-art open and closed-source VLMs, we identify significant performance gaps compared to human baselines, highlighting fundamental deficiencies in existing models. Extensive analysis reveals that VLMs struggle particularly with integrating multiple visual cues and maintaining temporal coherence. We further explore promising directions, such as leveraging 4D feature field reconstruction and targeted spatiotemporal supervised fine-tuning, demonstrating their effectiveness in enhancing spatiotemporal comprehension. Our work aims to encourage deeper exploration into improving VLMs' spatial and temporal grounding, paving the way towards more capable and reliable visual intelligence for dynamic environments.
title VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.02095