AnimatableDreamer: Text-Guided Non-rigid 3D Model Generation and Reconstruction with Canonical Score Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Xinzhou, Wang, Yikai, Ye, Junliang, Wang, Zhengyi, Sun, Fuchun, Liu, Pengkun, Wang, Ling, Sun, Kai, Wang, Xintong, He, Bin
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911817584869376
author Wang, Xinzhou
Wang, Yikai
Ye, Junliang
Wang, Zhengyi
Sun, Fuchun
Liu, Pengkun
Wang, Ling
Sun, Kai
Wang, Xintong
He, Bin
author_facet Wang, Xinzhou
Wang, Yikai
Ye, Junliang
Wang, Zhengyi
Sun, Fuchun
Liu, Pengkun
Wang, Ling
Sun, Kai
Wang, Xintong
He, Bin
contents Advances in 3D generation have facilitated sequential 3D model generation (a.k.a 4D generation), yet its application for animatable objects with large motion remains scarce. Our work proposes AnimatableDreamer, a text-to-4D generation framework capable of generating diverse categories of non-rigid objects on skeletons extracted from a monocular video. At its core, AnimatableDreamer is equipped with our novel optimization design dubbed Canonical Score Distillation (CSD), which lifts 2D diffusion for temporal consistent 4D generation. CSD, designed from a score gradient perspective, generates a canonical model with warp-robustness across different articulations. Notably, it also enhances the authenticity of bones and skinning by integrating inductive priors from a diffusion model. Furthermore, with multi-view distillation, CSD infers invisible regions, thereby improving the fidelity of monocular non-rigid reconstruction. Extensive experiments demonstrate the capability of our method in generating high-flexibility text-guided 3D models from the monocular video, while also showing improved reconstruction performance over existing non-rigid reconstruction methods.
format Preprint
id arxiv_https___arxiv_org_abs_2312_03795
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle AnimatableDreamer: Text-Guided Non-rigid 3D Model Generation and Reconstruction with Canonical Score Distillation
Wang, Xinzhou
Wang, Yikai
Ye, Junliang
Wang, Zhengyi
Sun, Fuchun
Liu, Pengkun
Wang, Ling
Sun, Kai
Wang, Xintong
He, Bin
Computer Vision and Pattern Recognition
I.4.5
Advances in 3D generation have facilitated sequential 3D model generation (a.k.a 4D generation), yet its application for animatable objects with large motion remains scarce. Our work proposes AnimatableDreamer, a text-to-4D generation framework capable of generating diverse categories of non-rigid objects on skeletons extracted from a monocular video. At its core, AnimatableDreamer is equipped with our novel optimization design dubbed Canonical Score Distillation (CSD), which lifts 2D diffusion for temporal consistent 4D generation. CSD, designed from a score gradient perspective, generates a canonical model with warp-robustness across different articulations. Notably, it also enhances the authenticity of bones and skinning by integrating inductive priors from a diffusion model. Furthermore, with multi-view distillation, CSD infers invisible regions, thereby improving the fidelity of monocular non-rigid reconstruction. Extensive experiments demonstrate the capability of our method in generating high-flexibility text-guided 3D models from the monocular video, while also showing improved reconstruction performance over existing non-rigid reconstruction methods.
title AnimatableDreamer: Text-Guided Non-rigid 3D Model Generation and Reconstruction with Canonical Score Distillation
topic Computer Vision and Pattern Recognition
I.4.5
url https://arxiv.org/abs/2312.03795