Saved in:
Bibliographic Details
Main Authors: Huang, Yuzhi, Wen, Kairun, Gao, Rongxin, Liu, Dongxuan, Lou, Yibin, Wu, Jie, Xu, Jing, Zhang, Jian, Yang, Zheng, Lin, Yunlong, Li, Chenxin, Pan, Panwang, Lu, Junbin, Jiang, Jingyan, Ding, Xinghao, Huang, Yue, Wang, Zhi
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.12746
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917338201194496
author Huang, Yuzhi
Wen, Kairun
Gao, Rongxin
Liu, Dongxuan
Lou, Yibin
Wu, Jie
Xu, Jing
Zhang, Jian
Yang, Zheng
Lin, Yunlong
Li, Chenxin
Pan, Panwang
Lu, Junbin
Jiang, Jingyan
Ding, Xinghao
Huang, Yue
Wang, Zhi
author_facet Huang, Yuzhi
Wen, Kairun
Gao, Rongxin
Liu, Dongxuan
Lou, Yibin
Wu, Jie
Xu, Jing
Zhang, Jian
Yang, Zheng
Lin, Yunlong
Li, Chenxin
Pan, Panwang
Lu, Junbin
Jiang, Jingyan
Ding, Xinghao
Huang, Yue
Wang, Zhi
contents Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLLMs) excel in static visual understanding, can they also be adept at "thinking in dynamics", i.e., perceive, track and reason about spatio-temporal dynamics in evolving scenes? To systematically assess their spatio-temporal reasoning and localized dynamics perception capabilities, we introduce Dyn-Bench, a large-scale benchmark built from diverse real-world and synthetic video datasets, enabling robust and scalable evaluation of spatio-temporal understanding. Through multi-stage filtering from massive 2D and 4D data sources, Dyn-Bench provides a high-quality collection of dynamic scenes, comprising 1k videos, 7k visual question answering (VQA) pairs, and 3k dynamic object grounding pairs. We probe general, spatial and region-level MLLMs to express how they think in dynamics both linguistically and visually, and find that existing models cannot simultaneously maintain strong performance in both spatio-temporal reasoning and dynamic object grounding, often producing inconsistent interpretations of motion and interaction. Notably, conventional prompting strategies (e.g., chain-of-thought or caption-based hints) provide limited improvement, whereas structured integration approaches, including Mask-Guided Fusion and Spatio-Temporal Textual Cognitive Map (ST-TCM), significantly enhance MLLMs' dynamics perception and spatio-temporal reasoning in the physical 4D world. Code and benchmark are available at https://dyn-bench.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12746
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
Huang, Yuzhi
Wen, Kairun
Gao, Rongxin
Liu, Dongxuan
Lou, Yibin
Wu, Jie
Xu, Jing
Zhang, Jian
Yang, Zheng
Lin, Yunlong
Li, Chenxin
Pan, Panwang
Lu, Junbin
Jiang, Jingyan
Ding, Xinghao
Huang, Yue
Wang, Zhi
Computer Vision and Pattern Recognition
Humans inhabit a physical 4D world where geometric structure and semantic content evolve over time, constituting a dynamic 4D reality (spatial with temporal dimension). While current Multimodal Large Language Models (MLLMs) excel in static visual understanding, can they also be adept at "thinking in dynamics", i.e., perceive, track and reason about spatio-temporal dynamics in evolving scenes? To systematically assess their spatio-temporal reasoning and localized dynamics perception capabilities, we introduce Dyn-Bench, a large-scale benchmark built from diverse real-world and synthetic video datasets, enabling robust and scalable evaluation of spatio-temporal understanding. Through multi-stage filtering from massive 2D and 4D data sources, Dyn-Bench provides a high-quality collection of dynamic scenes, comprising 1k videos, 7k visual question answering (VQA) pairs, and 3k dynamic object grounding pairs. We probe general, spatial and region-level MLLMs to express how they think in dynamics both linguistically and visually, and find that existing models cannot simultaneously maintain strong performance in both spatio-temporal reasoning and dynamic object grounding, often producing inconsistent interpretations of motion and interaction. Notably, conventional prompting strategies (e.g., chain-of-thought or caption-based hints) provide limited improvement, whereas structured integration approaches, including Mask-Guided Fusion and Spatio-Temporal Textual Cognitive Map (ST-TCM), significantly enhance MLLMs' dynamics perception and spatio-temporal reasoning in the physical 4D world. Code and benchmark are available at https://dyn-bench.github.io/.
title Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.12746