Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908008996405248 |
|---|---|
| author | Chen, Yabo Fang, Jiemin Huang, Yuyang Yi, Taoran Zhang, Xiaopeng Xie, Lingxi Wang, Xinggang Dai, Wenrui Xiong, Hongkai Tian, Qi |
| author_facet | Chen, Yabo Fang, Jiemin Huang, Yuyang Yi, Taoran Zhang, Xiaopeng Xie, Lingxi Wang, Xinggang Dai, Wenrui Xiong, Hongkai Tian, Qi |
| contents | Synthesizing multi-view 3D from one single image is a significant but challenging task. Zero-1-to-3 methods have achieved great success by lifting a 2D latent diffusion model to the 3D scope. The target view image is generated with a single-view source image and the camera pose as condition information. However, due to the high sparsity of the single input image, Zero-1-to-3 tends to produce geometry and appearance inconsistency across views, especially for complex objects. To tackle this issue, we propose to supply more condition information for the generation model but in a self-prompt way. A cascade framework is constructed with two Zero-1-to-3 models, named Cascade-Zero123, which progressively extract 3D information from the source image. Specifically, several nearby views are first generated by the first model and then fed into the second-stage model along with the source image as generation conditions. With amplified self-prompted condition images, our Cascade-Zero123 generates more consistent novel-view images than Zero-1-to-3. Experiment results demonstrate remarkable promotion, especially for various complex and challenging scenes, involving insects, humans, transparent objects, and stacked multiple objects etc. More demos and code are available at https://cascadezero123.github.io. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2312_04424 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views Chen, Yabo Fang, Jiemin Huang, Yuyang Yi, Taoran Zhang, Xiaopeng Xie, Lingxi Wang, Xinggang Dai, Wenrui Xiong, Hongkai Tian, Qi Computer Vision and Pattern Recognition Graphics Synthesizing multi-view 3D from one single image is a significant but challenging task. Zero-1-to-3 methods have achieved great success by lifting a 2D latent diffusion model to the 3D scope. The target view image is generated with a single-view source image and the camera pose as condition information. However, due to the high sparsity of the single input image, Zero-1-to-3 tends to produce geometry and appearance inconsistency across views, especially for complex objects. To tackle this issue, we propose to supply more condition information for the generation model but in a self-prompt way. A cascade framework is constructed with two Zero-1-to-3 models, named Cascade-Zero123, which progressively extract 3D information from the source image. Specifically, several nearby views are first generated by the first model and then fed into the second-stage model along with the source image as generation conditions. With amplified self-prompted condition images, our Cascade-Zero123 generates more consistent novel-view images than Zero-1-to-3. Experiment results demonstrate remarkable promotion, especially for various complex and challenging scenes, involving insects, humans, transparent objects, and stacked multiple objects etc. More demos and code are available at https://cascadezero123.github.io. |
| title | Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views |
| topic | Computer Vision and Pattern Recognition Graphics |
| url | https://arxiv.org/abs/2312.04424 |