Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yu, Qiao, Li, Xianzhi, Tang, Yuan, Han, Xu, Hu, Long, Hao, Yixue, Chen, Min
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909734124126208
author Yu, Qiao
Li, Xianzhi
Tang, Yuan
Han, Xu
Hu, Long
Hao, Yixue
Chen, Min
author_facet Yu, Qiao
Li, Xianzhi
Tang, Yuan
Han, Xu
Hu, Long
Hao, Yixue
Chen, Min
contents Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM's generated mesh ensures high clarity, discarding LRM's predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123's SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Code at: https://github.com/YuQiao0303/Fancy123
format Preprint
id arxiv_https___arxiv_org_abs_2411_16185
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
Yu, Qiao
Li, Xianzhi
Tang, Yuan
Han, Xu
Hu, Long
Hao, Yixue
Chen, Min
Computer Vision and Pattern Recognition
Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM's generated mesh ensures high clarity, discarding LRM's predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123's SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Code at: https://github.com/YuQiao0303/Fancy123
title Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.16185