Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866911572937408512 |
|---|---|
| author | Pan, Panwang Lin, Chenguo Zhao, Jingjing Li, Chenxin Lin, Yuchen Li, Haopeng Yan, Honglei Wen, Kairun Lin, Yunlong Yuan, Yixuan Mu, Yadong |
| author_facet | Pan, Panwang Lin, Chenguo Zhao, Jingjing Li, Chenxin Lin, Yuchen Li, Haopeng Yan, Honglei Wen, Kairun Lin, Yunlong Yuan, Yixuan Mu, Yadong |
| contents | We introduce Diff4Splat, a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets. Given a single input image, a camera trajectory, and an optional text prompt, Diff4Splat directly predicts a deformable 3D Gaussian field that encodes appearance, geometry, and motion, all in a single forward pass, without test-time optimization or post-hoc refinement. At the core of our framework lies a video latent transformer, which augments video diffusion models to jointly capture spatio-temporal dependencies and predict time-varying 3D Gaussian primitives. Training is guided by objectives on appearance fidelity, geometric accuracy, and motion consistency, enabling Diff4Splat to synthesize high-quality 4D scenes in 30 seconds. We demonstrate the effectiveness of Diff4Splat across video generation, novel view synthesis, and geometry extraction, where it matches or surpasses optimization-based methods for dynamic scene synthesis while being significantly more efficient. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_00503 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models Pan, Panwang Lin, Chenguo Zhao, Jingjing Li, Chenxin Lin, Yuchen Li, Haopeng Yan, Honglei Wen, Kairun Lin, Yunlong Yuan, Yixuan Mu, Yadong Computer Vision and Pattern Recognition We introduce Diff4Splat, a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from large-scale 4D datasets. Given a single input image, a camera trajectory, and an optional text prompt, Diff4Splat directly predicts a deformable 3D Gaussian field that encodes appearance, geometry, and motion, all in a single forward pass, without test-time optimization or post-hoc refinement. At the core of our framework lies a video latent transformer, which augments video diffusion models to jointly capture spatio-temporal dependencies and predict time-varying 3D Gaussian primitives. Training is guided by objectives on appearance fidelity, geometric accuracy, and motion consistency, enabling Diff4Splat to synthesize high-quality 4D scenes in 30 seconds. We demonstrate the effectiveness of Diff4Splat across video generation, novel view synthesis, and geometry extraction, where it matches or surpasses optimization-based methods for dynamic scene synthesis while being significantly more efficient. |
| title | Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2511.00503 |