Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917968574676992 |
|---|---|
| author | Liu, Tianqi Huang, Zihao Chen, Zhaoxi Wang, Guangcong Hu, Shoukang Shen, Liao Sun, Huiqiang Cao, Zhiguo Li, Wei Liu, Ziwei |
| author_facet | Liu, Tianqi Huang, Zihao Chen, Zhaoxi Wang, Guangcong Hu, Shoukang Shen, Liao Sun, Huiqiang Cao, Zhiguo Li, Wei Liu, Ziwei |
| contents | We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization ability due to the scarcity of 4D scene data. In contrast, our key insight is to distill pre-trained foundation models for consistent 4D scene representation, which offers promising advantages such as efficiency and generalizability. 1) To achieve this, we first animate the input image using image-to-video diffusion models followed by 4D geometric structure initialization. 2) To turn this coarse structure into spatial-temporal consistent multiview videos, we design an adaptive guidance mechanism with a point-guided denoising strategy for spatial consistency and a novel latent replacement strategy for temporal coherence. 3) To lift these generated observations into consistent 4D representation, we propose a modulation-based refinement to mitigate inconsistencies while fully leveraging the generated information. The resulting 4D representation enables real-time, controllable rendering, marking a significant advancement in single-image-based 4D scene generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_20785 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency Liu, Tianqi Huang, Zihao Chen, Zhaoxi Wang, Guangcong Hu, Shoukang Shen, Liao Sun, Huiqiang Cao, Zhiguo Li, Wei Liu, Ziwei Computer Vision and Pattern Recognition We present Free4D, a novel tuning-free framework for 4D scene generation from a single image. Existing methods either focus on object-level generation, making scene-level generation infeasible, or rely on large-scale multi-view video datasets for expensive training, with limited generalization ability due to the scarcity of 4D scene data. In contrast, our key insight is to distill pre-trained foundation models for consistent 4D scene representation, which offers promising advantages such as efficiency and generalizability. 1) To achieve this, we first animate the input image using image-to-video diffusion models followed by 4D geometric structure initialization. 2) To turn this coarse structure into spatial-temporal consistent multiview videos, we design an adaptive guidance mechanism with a point-guided denoising strategy for spatial consistency and a novel latent replacement strategy for temporal coherence. 3) To lift these generated observations into consistent 4D representation, we propose a modulation-based refinement to mitigate inconsistencies while fully leveraging the generated information. The resulting 4D representation enables real-time, controllable rendering, marking a significant advancement in single-image-based 4D scene generation. |
| title | Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2503.20785 |