HiFi-123: Towards High-fidelity One Image to 3D Content Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Wangbo, Yuan, Li, Cao, Yan-Pei, Gao, Xiangjun, Li, Xiaoyu, Hu, Wenbo, Quan, Long, Shan, Ying, Tian, Yonghong
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917719916412928
author Yu, Wangbo
Yuan, Li
Cao, Yan-Pei
Gao, Xiangjun
Li, Xiaoyu
Hu, Wenbo
Quan, Long
Shan, Ying
Tian, Yonghong
author_facet Yu, Wangbo
Yuan, Li
Cao, Yan-Pei
Gao, Xiangjun
Li, Xiaoyu
Hu, Wenbo
Quan, Long
Shan, Ying
Tian, Yonghong
contents Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123, a method designed for high-fidelity and multi-view consistent 3D generation. Our contributions are twofold: First, we propose a Reference-Guided Novel View Enhancement (RGNV) technique that significantly improves the fidelity of diffusion-based zero-shot novel view synthesis methods. Second, capitalizing on the RGNV, we present a novel Reference-Guided State Distillation (RGSD) loss. When incorporated into the optimization-based image-to-3D pipeline, our method significantly improves 3D generation quality, achieving state-of-the-art performance. Comprehensive evaluations demonstrate the effectiveness of our approach over existing methods, both qualitatively and quantitatively. Video results are available on the project page.
format Preprint
id arxiv_https___arxiv_org_abs_2310_06744
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle HiFi-123: Towards High-fidelity One Image to 3D Content Generation
Yu, Wangbo
Yuan, Li
Cao, Yan-Pei
Gao, Xiangjun
Li, Xiaoyu
Hu, Wenbo
Quan, Long
Shan, Ying
Tian, Yonghong
Computer Vision and Pattern Recognition
Recent advances in diffusion models have enabled 3D generation from a single image. However, current methods often produce suboptimal results for novel views, with blurred textures and deviations from the reference image, limiting their practical applications. In this paper, we introduce HiFi-123, a method designed for high-fidelity and multi-view consistent 3D generation. Our contributions are twofold: First, we propose a Reference-Guided Novel View Enhancement (RGNV) technique that significantly improves the fidelity of diffusion-based zero-shot novel view synthesis methods. Second, capitalizing on the RGNV, we present a novel Reference-Guided State Distillation (RGSD) loss. When incorporated into the optimization-based image-to-3D pipeline, our method significantly improves 3D generation quality, achieving state-of-the-art performance. Comprehensive evaluations demonstrate the effectiveness of our approach over existing methods, both qualitatively and quantitatively. Video results are available on the project page.
title HiFi-123: Towards High-fidelity One Image to 3D Content Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.06744