Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.07996 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908689647009792 |
|---|---|
| author | Kong, Lingdong Yang, Wesley Mei, Jianbiao Liu, Youquan Liang, Ao Zhu, Dekai Lu, Dongyue Yin, Wei Hu, Xiaotao Jia, Mingkai Deng, Junyuan Zhang, Kaiwen Wu, Yang Yan, Tianyi Gao, Shenyuan Wang, Song Li, Linfeng Pan, Liang Liu, Yong Zhu, Jianke Ooi, Wei Tsang Hoi, Steven C. H. Liu, Ziwei |
| author_facet | Kong, Lingdong Yang, Wesley Mei, Jianbiao Liu, Youquan Liang, Ao Zhu, Dekai Lu, Dongyue Yin, Wei Hu, Xiaotao Jia, Mingkai Deng, Junyuan Zhang, Kaiwen Wu, Yang Yan, Tianyi Gao, Shenyuan Wang, Song Li, Linfeng Pan, Liang Liu, Yong Zhu, Jianke Ooi, Wei Tsang Hoi, Steven C. H. Liu, Ziwei |
| contents | World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they overlook the rapidly growing body of work that leverages native 3D and 4D representations such as RGB-D imagery, occupancy grids, and LiDAR point clouds for large-scale scene modeling. At the same time, the absence of a standardized definition and taxonomy for ``world models'' has led to fragmented and sometimes inconsistent claims in the literature. This survey addresses these gaps by presenting the first comprehensive review explicitly dedicated to 3D and 4D world modeling and generation. We establish precise definitions, introduce a structured taxonomy spanning video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based (LiDARGen) approaches, and systematically summarize datasets and evaluation metrics tailored to 3D/4D settings. We further discuss practical applications, identify open challenges, and highlight promising research directions, aiming to provide a coherent and foundational reference for advancing the field. A systematic summary of existing literature is available at https://github.com/worldbench/awesome-3d-4d-world-models |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_07996 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | 3D and 4D World Modeling: A Survey Kong, Lingdong Yang, Wesley Mei, Jianbiao Liu, Youquan Liang, Ao Zhu, Dekai Lu, Dongyue Yin, Wei Hu, Xiaotao Jia, Mingkai Deng, Junyuan Zhang, Kaiwen Wu, Yang Yan, Tianyi Gao, Shenyuan Wang, Song Li, Linfeng Pan, Liang Liu, Yong Zhu, Jianke Ooi, Wei Tsang Hoi, Steven C. H. Liu, Ziwei Computer Vision and Pattern Recognition Robotics World modeling has become a cornerstone in AI research, enabling agents to understand, represent, and predict the dynamic environments they inhabit. While prior work largely emphasizes generative methods for 2D image and video data, they overlook the rapidly growing body of work that leverages native 3D and 4D representations such as RGB-D imagery, occupancy grids, and LiDAR point clouds for large-scale scene modeling. At the same time, the absence of a standardized definition and taxonomy for ``world models'' has led to fragmented and sometimes inconsistent claims in the literature. This survey addresses these gaps by presenting the first comprehensive review explicitly dedicated to 3D and 4D world modeling and generation. We establish precise definitions, introduce a structured taxonomy spanning video-based (VideoGen), occupancy-based (OccGen), and LiDAR-based (LiDARGen) approaches, and systematically summarize datasets and evaluation metrics tailored to 3D/4D settings. We further discuss practical applications, identify open challenges, and highlight promising research directions, aiming to provide a coherent and foundational reference for advancing the field. A systematic summary of existing literature is available at https://github.com/worldbench/awesome-3d-4d-world-models |
| title | 3D and 4D World Modeling: A Survey |
| topic | Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2509.07996 |