SCP: Spatial Causal Prediction in Video

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yanguang, Yang, Jie, Wu, Shengqiong, Hu, Shutong, Qiu, Hongbo, Wang, Yu, Zhang, Guijia, Ze, Tan Kai, Fei, Hao, Lin, Chia-Wen, Lee, Mong-Li, Hsu, Wynne
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913003606114304
author Zhao, Yanguang
Yang, Jie
Wu, Shengqiong
Hu, Shutong
Qiu, Hongbo
Wang, Yu
Zhang, Guijia
Ze, Tan Kai
Fei, Hao
Lin, Chia-Wen
Lee, Mong-Li
Hsu, Wynne
author_facet Zhao, Yanguang
Yang, Jie
Wu, Shengqiong
Hu, Shutong
Qiu, Hongbo
Wang, Yu
Zhang, Guijia
Ze, Tan Kai
Fei, Hao
Lin, Chia-Wen
Lee, Mong-Li
Hsu, Wynne
contents Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as autonomous driving and robotics. Existing studies, however, primarily assess models on visible spatio-temporal understanding, overlooking their ability to infer unseen past or future spatial states. In this work, we introduce Spatial Causal Prediction (SCP), a new task paradigm that challenges models to reason beyond observation and predict spatial causal outcomes. We further construct SCP-Bench, a benchmark comprising 2,500 QA pairs across 1,181 videos spanning diverse viewpoints, scenes, and causal directions, to support systematic evaluation. Through comprehensive experiments on {23} state-of-the-art models, we reveal substantial gaps between human and model performance, limited temporal extrapolation, and weak causal grounding. We further analyze key factors influencing performance and propose perception-enhancement and reasoning-guided strategies toward advancing spatial causal intelligence. The project page is https://guangstrip.github.io/SCP-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2603_03944
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SCP: Spatial Causal Prediction in Video
Zhao, Yanguang
Yang, Jie
Wu, Shengqiong
Hu, Shutong
Qiu, Hongbo
Wang, Yu
Zhang, Guijia
Ze, Tan Kai
Fei, Hao
Lin, Chia-Wen
Lee, Mong-Li
Hsu, Wynne
Computer Vision and Pattern Recognition
Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as autonomous driving and robotics. Existing studies, however, primarily assess models on visible spatio-temporal understanding, overlooking their ability to infer unseen past or future spatial states. In this work, we introduce Spatial Causal Prediction (SCP), a new task paradigm that challenges models to reason beyond observation and predict spatial causal outcomes. We further construct SCP-Bench, a benchmark comprising 2,500 QA pairs across 1,181 videos spanning diverse viewpoints, scenes, and causal directions, to support systematic evaluation. Through comprehensive experiments on {23} state-of-the-art models, we reveal substantial gaps between human and model performance, limited temporal extrapolation, and weak causal grounding. We further analyze key factors influencing performance and propose perception-enhancement and reasoning-guided strategies toward advancing spatial causal intelligence. The project page is https://guangstrip.github.io/SCP-Bench.
title SCP: Spatial Causal Prediction in Video
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.03944