Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912641451032576 |
|---|---|
| author | Cai, Yuanhao Zhang, He Zhang, Kai Liang, Yixun Ren, Mengwei Luan, Fujun Liu, Qing Kim, Soo Ye Zhang, Jianming Zhang, Zhifei Zhou, Yuqian Zhang, Yulun Yang, Xiaokang Lin, Zhe Yuille, Alan |
| author_facet | Cai, Yuanhao Zhang, He Zhang, Kai Liang, Yixun Ren, Mengwei Luan, Fujun Liu, Qing Kim, Soo Ye Zhang, Jianming Zhang, Zhifei Zhou, Yuqian Zhang, Yulun Yang, Xiaokang Lin, Zhe Yuille, Alan |
| contents | Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model, DiffusionGS, for object generation and scene reconstruction from a single view. DiffusionGS directly outputs 3D Gaussian point clouds at each timestep to enforce view consistency and allow the model to generate robustly given prompt views of any directions, beyond object-centric inputs. Plus, to improve the capability and generality of DiffusionGS, we scale up 3D training data by developing a scene-object mixed training strategy. Experiments show that DiffusionGS yields improvements of 2.20 dB/23.25 and 1.34 dB/19.16 in PSNR/FID for objects and scenes than the state-of-the-art methods, without depth estimator. Plus, our method enjoys over 5$\times$ faster speed ($\sim$6s on an A100 GPU). Our Project page at https://caiyuanhao1998.github.io/project/DiffusionGS/ shows the video and interactive results. The code and models are publicly available at https://github.com/caiyuanhao1998/Open-DiffusionGS |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_14384 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction Cai, Yuanhao Zhang, He Zhang, Kai Liang, Yixun Ren, Mengwei Luan, Fujun Liu, Qing Kim, Soo Ye Zhang, Jianming Zhang, Zhifei Zhou, Yuqian Zhang, Yulun Yang, Xiaokang Lin, Zhe Yuille, Alan Computer Vision and Pattern Recognition Graphics Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model, DiffusionGS, for object generation and scene reconstruction from a single view. DiffusionGS directly outputs 3D Gaussian point clouds at each timestep to enforce view consistency and allow the model to generate robustly given prompt views of any directions, beyond object-centric inputs. Plus, to improve the capability and generality of DiffusionGS, we scale up 3D training data by developing a scene-object mixed training strategy. Experiments show that DiffusionGS yields improvements of 2.20 dB/23.25 and 1.34 dB/19.16 in PSNR/FID for objects and scenes than the state-of-the-art methods, without depth estimator. Plus, our method enjoys over 5$\times$ faster speed ($\sim$6s on an A100 GPU). Our Project page at https://caiyuanhao1998.github.io/project/DiffusionGS/ shows the video and interactive results. The code and models are publicly available at https://github.com/caiyuanhao1998/Open-DiffusionGS |
| title | Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction |
| topic | Computer Vision and Pattern Recognition Graphics |
| url | https://arxiv.org/abs/2411.14384 |