Reconstructing 4D Spatial Intelligence: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Yukang, Lu, Jiahao, Huang, Zhisheng, Shen, Zhuowen, Zhao, Chengfeng, Hong, Fangzhou, Chen, Zhaoxi, Li, Xin, Wang, Wenping, Liu, Yuan, Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911087817916416
author Cao, Yukang
Lu, Jiahao
Huang, Zhisheng
Shen, Zhuowen
Zhao, Chengfeng
Hong, Fangzhou
Chen, Zhaoxi
Li, Xin
Wang, Wenping
Liu, Yuan
Liu, Ziwei
author_facet Cao, Yukang
Lu, Jiahao
Huang, Zhisheng
Shen, Zhuowen
Zhao, Chengfeng
Hong, Fangzhou
Chen, Zhaoxi
Li, Xin
Wang, Wenping
Liu, Yuan
Liu, Ziwei
contents Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21045
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reconstructing 4D Spatial Intelligence: A Survey
Cao, Yukang
Lu, Jiahao
Huang, Zhisheng
Shen, Zhuowen
Zhao, Chengfeng
Hong, Fangzhou
Chen, Zhaoxi
Li, Xin
Wang, Wenping
Liu, Yuan
Liu, Ziwei
Computer Vision and Pattern Recognition
Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, where the focus is often on reconstructing fundamental visual elements, to embodied AI, which emphasizes interaction modeling and physical realism. Fueled by rapid advances in 3D representations and deep learning architectures, the field has evolved quickly, outpacing the scope of previous surveys. Additionally, existing surveys rarely offer a comprehensive analysis of the hierarchical structure of 4D scene reconstruction. To address this gap, we present a new perspective that organizes existing methods into five progressive levels of 4D spatial intelligence: (1) Level 1 -- reconstruction of low-level 3D attributes (e.g., depth, pose, and point maps); (2) Level 2 -- reconstruction of 3D scene components (e.g., objects, humans, structures); (3) Level 3 -- reconstruction of 4D dynamic scenes; (4) Level 4 -- modeling of interactions among scene components; and (5) Level 5 -- incorporation of physical laws and constraints. We conclude the survey by discussing the key challenges at each level and highlighting promising directions for advancing toward even richer levels of 4D spatial intelligence. To track ongoing developments, we maintain an up-to-date project page: https://github.com/yukangcao/Awesome-4D-Spatial-Intelligence.
title Reconstructing 4D Spatial Intelligence: A Survey
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.21045