Saved in:
Bibliographic Details
Main Authors: Kong, Lingdong, Yang, Wesley, Mei, Jianbiao, Liu, Youquan, Liang, Ao, Zhu, Dekai, Lu, Dongyue, Yin, Wei, Hu, Xiaotao, Jia, Mingkai, Deng, Junyuan, Zhang, Kaiwen, Wu, Yang, Yan, Tianyi, Gao, Shenyuan, Wang, Song, Li, Linfeng, Pan, Liang, Liu, Yong, Zhu, Jianke, Ooi, Wei Tsang, Hoi, Steven C. H., Liu, Ziwei
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.07996
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908689647009792
author Kong, Lingdong
Yang, Wesley
Mei, Jianbiao
Liu, Youquan
Liang, Ao
Zhu, Dekai
Lu, Dongyue
Yin, Wei
Hu, Xiaotao
Jia, Mingkai
Deng, Junyuan
Zhang, Kaiwen
Wu, Yang
Yan, Tianyi
Gao, Shenyuan
Wang, Song
Li, Linfeng
Pan, Liang
Liu, Yong
Zhu, Jianke
Ooi, Wei Tsang
Hoi, Steven C. H.
Liu, Ziwei
author_facet Kong, Lingdong
Yang, Wesley
Mei, Jianbiao
Liu, Youquan
Liang, Ao
Zhu, Dekai
Lu, Dongyue
Yin, Wei
Hu, Xiaotao
Jia, Mingkai
Deng, Junyuan
Zhang, Kaiwen
Wu, Yang
Yan, Tianyi
Gao, Shenyuan
Wang, Song
Li, Linfeng
Pan, Liang
Liu, Yong
Zhu, Jianke
Ooi, Wei Tsang
Hoi, Steven C. H.
Liu, Ziwei
contents World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they overlook the rapidly growing body of work that leverages native 3D and 4D representations such as RGB-D imagery, occupancy grids, and LiDAR point clouds for large-scale scene modeling. At the same time, the absence of a standardized definition and taxonomy for ``world models'' has led to fragmented and sometimes inconsistent claims in the literature. This survey addresses these gaps by presenting the first comprehensive review explicitly dedicated to 3D and 4D world modeling and generation. We establish precise definitions, introduce a structured taxonomy spanning video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based (LiDARGen) approaches, and systematically summarize datasets and evaluation metrics tailored to 3D/4D settings. We further discuss practical applications, identify open challenges, and highlight promising research directions, aiming to provide a coherent and foundational reference for advancing the field. A systematic summary of existing literature is available at https://github.com/worldbench/awesome-3d-4d-world-models
format Preprint
id arxiv_https___arxiv_org_abs_2509_07996
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle 3D and 4D World Modeling: A Survey
Kong, Lingdong
Yang, Wesley
Mei, Jianbiao
Liu, Youquan
Liang, Ao
Zhu, Dekai
Lu, Dongyue
Yin, Wei
Hu, Xiaotao
Jia, Mingkai
Deng, Junyuan
Zhang, Kaiwen
Wu, Yang
Yan, Tianyi
Gao, Shenyuan
Wang, Song
Li, Linfeng
Pan, Liang
Liu, Yong
Zhu, Jianke
Ooi, Wei Tsang
Hoi, Steven C. H.
Liu, Ziwei
Computer Vision and Pattern Recognition
Robotics
World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they overlook the rapidly growing body of work that leverages native 3D and 4D representations such as RGB-D imagery, occupancy grids, and LiDAR point clouds for large-scale scene modeling. At the same time, the absence of a standardized definition and taxonomy for ``world models'' has led to fragmented and sometimes inconsistent claims in the literature. This survey addresses these gaps by presenting the first comprehensive review explicitly dedicated to 3D and 4D world modeling and generation. We establish precise definitions, introduce a structured taxonomy spanning video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based (LiDARGen) approaches, and systematically summarize datasets and evaluation metrics tailored to 3D/4D settings. We further discuss practical applications, identify open challenges, and highlight promising research directions, aiming to provide a coherent and foundational reference for advancing the field. A systematic summary of existing literature is available at https://github.com/worldbench/awesome-3d-4d-world-models
title 3D and 4D World Modeling: A Survey
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2509.07996