DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Jiazhe, Ding, Yikang, Chen, Xiwu, Chen, Shuo, Li, Bohan, Zou, Yingshuang, Lyu, Xiaoyang, Tan, Feiyang, Qi, Xiaojuan, Li, Zhiheng, Zhao, Hao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
by: Zou, Yingshuang, et al.
Published: (2025)
by: Zou, Yingshuang, et al.
Published: (2025)
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
M${^2}$Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
by: Zou, Yingshuang, et al.
Published: (2024)
by: Zou, Yingshuang, et al.
Published: (2024)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)
by: Zhou, Xin, et al.
Published: (2026)
4D Driving Scene Generation With Stereo Forcing
by: Lu, Hao, et al.
Published: (2025)
by: Lu, Hao, et al.
Published: (2025)
Total-Decom: Decomposed 3D Scene Reconstruction with Minimal Interaction
by: Lyu, Xiaoyang, et al.
Published: (2024)
by: Lyu, Xiaoyang, et al.
Published: (2024)
ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
by: Wang, Haonan, et al.
Published: (2026)
by: Wang, Haonan, et al.
Published: (2026)
UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
by: Shi, Chen, et al.
Published: (2025)
by: Shi, Chen, et al.
Published: (2025)
ST-VLA: Enabling 4D-Aware Spatiotemporal Understanding for General Robot Manipulation
by: Wu, You, et al.
Published: (2026)
by: Wu, You, et al.
Published: (2026)
DO3D: Self-supervised Learning of Decomposed Object-aware 3D Motion and Depth from Monocular Videos
by: Wu, Xiuzhe, et al.
Published: (2024)
by: Wu, Xiuzhe, et al.
Published: (2024)
The Devil is in the Edges: Monocular Depth Estimation with Edge-aware Consistency Fusion
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
by: Chen, Xiaoxue, et al.
Published: (2025)
by: Chen, Xiaoxue, et al.
Published: (2025)
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation
by: Zhao, Guosheng, et al.
Published: (2024)
by: Zhao, Guosheng, et al.
Published: (2024)
4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
by: Xie, Siyi, et al.
Published: (2025)
by: Xie, Siyi, et al.
Published: (2025)
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
by: Li, Yiming, et al.
Published: (2023)
by: Li, Yiming, et al.
Published: (2023)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
by: Hu, Yu, et al.
Published: (2025)
by: Hu, Yu, et al.
Published: (2025)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving
by: He, Zhuolin, et al.
Published: (2026)
by: He, Zhuolin, et al.
Published: (2026)
EAG3R: Event-Augmented 3D Geometry Estimation for Dynamic and Extreme-Lighting Scenes
by: Wu, Xiaoshan, et al.
Published: (2025)
by: Wu, Xiaoshan, et al.
Published: (2025)
DiMeR: Disentangled Mesh Reconstruction Model
by: Jiang, Lutao, et al.
Published: (2025)
by: Jiang, Lutao, et al.
Published: (2025)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
ART3D: 3D Gaussian Splatting for Text-Guided Artistic Scenes Generation
by: Li, Pengzhi, et al.
Published: (2024)
by: Li, Pengzhi, et al.
Published: (2024)
DriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving
by: Min, Chen, et al.
Published: (2024)
by: Min, Chen, et al.
Published: (2024)
ORV: 4D Occupancy-centric Robot Video Generation
by: Yang, Xiuyu, et al.
Published: (2025)
by: Yang, Xiuyu, et al.
Published: (2025)
Disentangling Instance and Scene Contexts for 3D Semantic Scene Completion
by: Liu, Enyu, et al.
Published: (2025)
by: Liu, Enyu, et al.
Published: (2025)
4DRadar-GS: Self-Supervised Dynamic Driving Scene Reconstruction with 4D Radar
by: Tang, Xiao, et al.
Published: (2025)
by: Tang, Xiao, et al.
Published: (2025)
InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
by: Liu, Hongyuan, et al.
Published: (2025)
by: Liu, Hongyuan, et al.
Published: (2025)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagation
by: Chen, Shifeng, et al.
Published: (2026)
by: Chen, Shifeng, et al.
Published: (2026)
TranSplat: Generalizable 3D Gaussian Splatting from Sparse Multi-View Images with Transformers
by: Zhang, Chuanrui, et al.
Published: (2024)
by: Zhang, Chuanrui, et al.
Published: (2024)
VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes
by: Zhang, Yikang, et al.
Published: (2025)
by: Zhang, Yikang, et al.
Published: (2025)
Ground4D: Spatially-Grounded Feedforward 4D Reconstruction for Unstructured Off-Road Scenes
by: Wang, Shuo, et al.
Published: (2026)
by: Wang, Shuo, et al.
Published: (2026)
4D-CS: Exploiting Cluster Prior for 4D Spatio-Temporal LiDAR Semantic Segmentation
by: Zhong, Jiexi, et al.
Published: (2025)
by: Zhong, Jiexi, et al.
Published: (2025)
WideRange4D: Enabling High-Quality 4D Reconstruction with Wide-Range Movements and Scenes
by: Yang, Ling, et al.
Published: (2025)
by: Yang, Ling, et al.
Published: (2025)
4DRaL: Bridging 4D Radar with LiDAR for Place Recognition using Knowledge Distillation
by: Huang, Ningyuan, et al.
Published: (2026)
by: Huang, Ningyuan, et al.
Published: (2026)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
by: Xu, Mutian, et al.
Published: (2026)
by: Xu, Mutian, et al.
Published: (2026)
4DEquine: Disentangling Motion and Appearance for 4D Equine Reconstruction from Monocular Video
by: Lyu, Jin, et al.
Published: (2026)
by: Lyu, Jin, et al.
Published: (2026)
Similar Items
-
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
by: Zou, Yingshuang, et al.
Published: (2025) -
UniScene: Unified Occupancy-centric Driving Scene Generation
by: Li, Bohan, et al.
Published: (2024) -
M${^2}$Depth: Self-supervised Two-Frame Multi-camera Metric Depth Estimation
by: Zou, Yingshuang, et al.
Published: (2024) -
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2025) -
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)