DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Haoran, Shi, Haolin, Zhang, Wenli, Wu, Wenjun, Liao, Yong, Wang, Lin, Lee, Lik-hang, Zhou, Pengyuan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914876783329280
author Li, Haoran
Shi, Haolin
Zhang, Wenli
Wu, Wenjun
Liao, Yong
Wang, Lin
Lee, Lik-hang
Zhou, Pengyuan
author_facet Li, Haoran
Shi, Haolin
Zhang, Wenli
Wu, Wenjun
Liao, Yong
Wang, Lin
Lee, Lik-hang
Zhou, Pengyuan
contents Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper, we propose DreamScene, a 3D Gaussian-based novel text-to-3D scene generation framework, to tackle the aforementioned three challenges mainly via two strategies. First, DreamScene employs Formation Pattern Sampling (FPS), a multi-timestep sampling strategy guided by the formation patterns of 3D objects, to form fast, semantically rich, and high-quality representations. FPS uses 3D Gaussian filtering for optimization stability, and leverages reconstruction techniques to generate plausible textures. Second, DreamScene employs a progressive three-stage camera sampling strategy, specifically designed for both indoor and outdoor settings, to effectively ensure object-environment integration and scene-wide 3D consistency. Last, DreamScene enhances scene editing flexibility by integrating objects and environments, enabling targeted adjustments. Extensive experiments validate DreamScene's superiority over current state-of-the-art techniques, heralding its wide-ranging potential for diverse applications. Code and demos will be released at https://dreamscene-project.github.io .
format Preprint
id arxiv_https___arxiv_org_abs_2404_03575
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
Li, Haoran
Shi, Haolin
Zhang, Wenli
Wu, Wenjun
Liao, Yong
Wang, Lin
Lee, Lik-hang
Zhou, Pengyuan
Computer Vision and Pattern Recognition
Text-to-3D scene generation holds immense potential for the gaming, film, and architecture sectors. Despite significant progress, existing methods struggle with maintaining high quality, consistency, and editing flexibility. In this paper, we propose DreamScene, a 3D Gaussian-based novel text-to-3D scene generation framework, to tackle the aforementioned three challenges mainly via two strategies. First, DreamScene employs Formation Pattern Sampling (FPS), a multi-timestep sampling strategy guided by the formation patterns of 3D objects, to form fast, semantically rich, and high-quality representations. FPS uses 3D Gaussian filtering for optimization stability, and leverages reconstruction techniques to generate plausible textures. Second, DreamScene employs a progressive three-stage camera sampling strategy, specifically designed for both indoor and outdoor settings, to effectively ensure object-environment integration and scene-wide 3D consistency. Last, DreamScene enhances scene editing flexibility by integrating objects and environments, enabling targeted adjustments. Extensive experiments validate DreamScene's superiority over current state-of-the-art techniques, heralding its wide-ranging potential for diverse applications. Code and demos will be released at https://dreamscene-project.github.io .
title DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.03575