Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zhenyu, Zhang, Junwu, Cheng, Xinhua, Yu, Wangbo, Feng, Chaoran, Pang, Yatian, Lin, Bin, Yuan, Li
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929439765430272
author Tang, Zhenyu
Zhang, Junwu
Cheng, Xinhua
Yu, Wangbo
Feng, Chaoran
Pang, Yatian
Lin, Bin
Yuan, Li
author_facet Tang, Zhenyu
Zhang, Junwu
Cheng, Xinhua
Yu, Wangbo
Feng, Chaoran
Pang, Yatian
Lin, Bin
Yuan, Li
contents Recent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content.However, multi-view diffusion models often produce low-quality and inconsistent images, adversely affecting the quality of the final 3D reconstruction. To address this issue, we propose a unified 3D generation framework called Cycle3D, which cyclically utilizes a 2D diffusion-based generation module and a feed-forward 3D reconstruction module during the multi-step diffusion process. Concretely, 2D diffusion model is applied for generating high-quality texture, and the reconstruction model guarantees multi-view consistency.Moreover, 2D diffusion model can further control the generated content and inject reference-view information for unseen views, thereby enhancing the diversity and texture consistency of 3D generation during the denoising process. Extensive experiments demonstrate the superior ability of our method to create 3D content with high-quality and consistency compared with state-of-the-art baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19548
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
Tang, Zhenyu
Zhang, Junwu
Cheng, Xinhua
Yu, Wangbo
Feng, Chaoran
Pang, Yatian
Lin, Bin
Yuan, Li
Computer Vision and Pattern Recognition
Recent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content.However, multi-view diffusion models often produce low-quality and inconsistent images, adversely affecting the quality of the final 3D reconstruction. To address this issue, we propose a unified 3D generation framework called Cycle3D, which cyclically utilizes a 2D diffusion-based generation module and a feed-forward 3D reconstruction module during the multi-step diffusion process. Concretely, 2D diffusion model is applied for generating high-quality texture, and the reconstruction model guarantees multi-view consistency.Moreover, 2D diffusion model can further control the generated content and inject reference-view information for unseen views, thereby enhancing the diversity and texture consistency of 3D generation during the denoising process. Extensive experiments demonstrate the superior ability of our method to create 3D content with high-quality and consistency compared with state-of-the-art baselines.
title Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.19548