DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xiaoxue, Xiong, Ziyi, Chen, Yuantao, Li, Gen, Wang, Nan, Luo, Hongcheng, Chen, Long, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Li, Hongyang, Zhang, Ya-Qin, Zhao, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917119655936000
author Chen, Xiaoxue
Xiong, Ziyi
Chen, Yuantao
Li, Gen
Wang, Nan
Luo, Hongcheng
Chen, Long
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Li, Hongyang
Zhang, Ya-Qin
Zhao, Hao
author_facet Chen, Xiaoxue
Xiong, Ziyi
Chen, Yuantao
Li, Gen
Wang, Nan
Luo, Hongcheng
Chen, Long
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Li, Hongyang
Zhang, Ya-Qin
Zhao, Hao
contents Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows, making them slow and impractical. We revisit this problem from a feedforward perspective and introduce \textbf{Driving Gaussian Grounded Transformer (DGGT)}, a unified framework for pose-free dynamic scene reconstruction. We note that the existing formulations, treating camera pose as a required input, limit flexibility and scalability. Instead, we reformulate pose as an output of the model, enabling reconstruction directly from sparse, unposed images and supporting an arbitrary number of views for long sequences. Our approach jointly predicts per-frame 3D Gaussian maps and camera parameters, disentangles dynamics with a lightweight dynamic head, and preserves temporal consistency with a lifespan head that modulates visibility over time. A diffusion-based rendering refinement further reduces motion/interpolation artifacts and improves novel-view quality under sparse inputs. The result is a single-pass, pose-free algorithm that achieves state-of-the-art performance and speed. Trained and evaluated on large-scale driving benchmarks (Waymo, nuScenes, Argoverse2), our method outperforms prior work both when trained on each dataset and in zero-shot transfer across datasets, and it scales well as the number of input frames increases.
format Preprint
id arxiv_https___arxiv_org_abs_2512_03004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
Chen, Xiaoxue
Xiong, Ziyi
Chen, Yuantao
Li, Gen
Wang, Nan
Luo, Hongcheng
Chen, Long
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Li, Hongyang
Zhang, Ya-Qin
Zhao, Hao
Computer Vision and Pattern Recognition
Autonomous driving needs fast, scalable 4D reconstruction and re-simulation for training and evaluation, yet most methods for dynamic driving scenes still rely on per-scene optimization, known camera calibration, or short frame windows, making them slow and impractical. We revisit this problem from a feedforward perspective and introduce \textbf{Driving Gaussian Grounded Transformer (DGGT)}, a unified framework for pose-free dynamic scene reconstruction. We note that the existing formulations, treating camera pose as a required input, limit flexibility and scalability. Instead, we reformulate pose as an output of the model, enabling reconstruction directly from sparse, unposed images and supporting an arbitrary number of views for long sequences. Our approach jointly predicts per-frame 3D Gaussian maps and camera parameters, disentangles dynamics with a lightweight dynamic head, and preserves temporal consistency with a lifespan head that modulates visibility over time. A diffusion-based rendering refinement further reduces motion/interpolation artifacts and improves novel-view quality under sparse inputs. The result is a single-pass, pose-free algorithm that achieves state-of-the-art performance and speed. Trained and evaluated on large-scale driving benchmarks (Waymo, nuScenes, Argoverse2), our method outperforms prior work both when trained on each dataset and in zero-shot transfer across datasets, and it scales well as the number of input frames increases.
title DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.03004