Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cai, Yuanhao, Zhang, He, Zhang, Kai, Liang, Yixun, Ren, Mengwei, Luan, Fujun, Liu, Qing, Kim, Soo Ye, Zhang, Jianming, Zhang, Zhifei, Zhou, Yuqian, Zhang, Yulun, Yang, Xiaokang, Lin, Zhe, Yuille, Alan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912641451032576
author Cai, Yuanhao
Zhang, He
Zhang, Kai
Liang, Yixun
Ren, Mengwei
Luan, Fujun
Liu, Qing
Kim, Soo Ye
Zhang, Jianming
Zhang, Zhifei
Zhou, Yuqian
Zhang, Yulun
Yang, Xiaokang
Lin, Zhe
Yuille, Alan
author_facet Cai, Yuanhao
Zhang, He
Zhang, Kai
Liang, Yixun
Ren, Mengwei
Luan, Fujun
Liu, Qing
Kim, Soo Ye
Zhang, Jianming
Zhang, Zhifei
Zhou, Yuqian
Zhang, Yulun
Yang, Xiaokang
Lin, Zhe
Yuille, Alan
contents Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model, DiffusionGS, for object generation and scene reconstruction from a single view. DiffusionGS directly outputs 3D Gaussian point clouds at each timestep to enforce view consistency and allow the model to generate robustly given prompt views of any directions, beyond object-centric inputs. Plus, to improve the capability and generality of DiffusionGS, we scale up 3D training data by developing a scene-object mixed training strategy. Experiments show that DiffusionGS yields improvements of 2.20 dB/23.25 and 1.34 dB/19.16 in PSNR/FID for objects and scenes than the state-of-the-art methods, without depth estimator. Plus, our method enjoys over 5$\times$ faster speed ($\sim$6s on an A100 GPU). Our Project page at https://caiyuanhao1998.github.io/project/DiffusionGS/ shows the video and interactive results. The code and models are publicly available at https://github.com/caiyuanhao1998/Open-DiffusionGS
format Preprint
id arxiv_https___arxiv_org_abs_2411_14384
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
Cai, Yuanhao
Zhang, He
Zhang, Kai
Liang, Yixun
Ren, Mengwei
Luan, Fujun
Liu, Qing
Kim, Soo Ye
Zhang, Jianming
Zhang, Zhifei
Zhou, Yuqian
Zhang, Yulun
Yang, Xiaokang
Lin, Zhe
Yuille, Alan
Computer Vision and Pattern Recognition
Graphics
Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model, DiffusionGS, for object generation and scene reconstruction from a single view. DiffusionGS directly outputs 3D Gaussian point clouds at each timestep to enforce view consistency and allow the model to generate robustly given prompt views of any directions, beyond object-centric inputs. Plus, to improve the capability and generality of DiffusionGS, we scale up 3D training data by developing a scene-object mixed training strategy. Experiments show that DiffusionGS yields improvements of 2.20 dB/23.25 and 1.34 dB/19.16 in PSNR/FID for objects and scenes than the state-of-the-art methods, without depth estimator. Plus, our method enjoys over 5$\times$ faster speed ($\sim$6s on an A100 GPU). Our Project page at https://caiyuanhao1998.github.io/project/DiffusionGS/ shows the video and interactive results. The code and models are publicly available at https://github.com/caiyuanhao1998/Open-DiffusionGS
title Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
topic Computer Vision and Pattern Recognition
Graphics
url https://arxiv.org/abs/2411.14384