Retrieval-Augmented Score Distillation for Text-to-3D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Seo, Junyoung, Hong, Susung, Jang, Wooseok, Kim, Inès Hyeonsu, Kwak, Minseop, Lee, Doyup, Kim, Seungryong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914780140273664
author Seo, Junyoung
Hong, Susung
Jang, Wooseok
Kim, Inès Hyeonsu
Kwak, Minseop
Lee, Doyup
Kim, Seungryong
author_facet Seo, Junyoung
Hong, Susung
Jang, Wooseok
Kim, Inès Hyeonsu
Kwak, Minseop
Lee, Doyup
Kim, Seungryong
contents Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have been released, fine-tuning the diffusion model on the multi-view datasets becomes a mainstream to solve the 3D inconsistency problem. However, it has confronted with fundamental difficulties regarding the limited quality and diversity of 3D data, compared with 2D data. To sidestep these trade-offs, we explore a retrieval-augmented approach tailored for score distillation, dubbed ReDream. We postulate that both expressiveness of 2D diffusion models and geometric consistency of 3D assets can be fully leveraged by employing the semantically relevant assets directly within the optimization process. To this end, we introduce novel framework for retrieval-based quality enhancement in text-to-3D generation. We leverage the retrieved asset to incorporate its geometric prior in the variational objective and adapt the diffusion model's 2D prior toward view consistency, achieving drastic improvements in both geometry and fidelity of generated scenes. We conduct extensive experiments to demonstrate that ReDream exhibits superior quality with increased geometric consistency. Project page is available at https://ku-cvlab.github.io/ReDream/.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02972
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Retrieval-Augmented Score Distillation for Text-to-3D Generation
Seo, Junyoung
Hong, Susung
Jang, Wooseok
Kim, Inès Hyeonsu
Kwak, Minseop
Lee, Doyup
Kim, Seungryong
Computer Vision and Pattern Recognition
Machine Learning
Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-view datasets have been released, fine-tuning the diffusion model on the multi-view datasets becomes a mainstream to solve the 3D inconsistency problem. However, it has confronted with fundamental difficulties regarding the limited quality and diversity of 3D data, compared with 2D data. To sidestep these trade-offs, we explore a retrieval-augmented approach tailored for score distillation, dubbed ReDream. We postulate that both expressiveness of 2D diffusion models and geometric consistency of 3D assets can be fully leveraged by employing the semantically relevant assets directly within the optimization process. To this end, we introduce novel framework for retrieval-based quality enhancement in text-to-3D generation. We leverage the retrieved asset to incorporate its geometric prior in the variational objective and adapt the diffusion model's 2D prior toward view consistency, achieving drastic improvements in both geometry and fidelity of generated scenes. We conduct extensive experiments to demonstrate that ReDream exhibits superior quality with increased geometric consistency. Project page is available at https://ku-cvlab.github.io/ReDream/.
title Retrieval-Augmented Score Distillation for Text-to-3D Generation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2402.02972