HeadArtist: Text-conditioned 3D Head Generation with Self Score Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Hongyu, Wang, Xuan, Wan, Ziyu, Shen, Yujun, Song, Yibing, Liao, Jing, Chen, Qifeng
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909194495459328
author Liu, Hongyu
Wang, Xuan
Wan, Ziyu
Shen, Yujun
Song, Yibing
Liao, Jing
Chen, Qifeng
author_facet Liu, Hongyu
Wang, Xuan
Wan, Ziyu
Shen, Yujun
Song, Yibing
Liao, Jing
Chen, Qifeng
contents This work presents HeadArtist for 3D head generation from text descriptions. With a landmark-guided ControlNet serving as the generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the supervision of the prior distillation itself. We call such a process self score distillation (SSD). In detail, given a sampled camera pose, we first render an image and its corresponding landmarks from the head model, and add some particular level of noise onto the image. The noisy image, landmarks, and text condition are then fed into the frozen ControlNet twice for noise prediction. Two different classifier-free guidance (CFG) weights are applied during these two predictions, and the prediction difference offers a direction on how the rendered image can better match the text of interest. Experimental results suggest that our approach delivers high-quality 3D head sculptures with adequate geometry and photorealistic appearance, significantly outperforming state-ofthe-art methods. We also show that the same pipeline well supports editing the generated heads, including both geometry deformation and appearance change.
format Preprint
id arxiv_https___arxiv_org_abs_2312_07539
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle HeadArtist: Text-conditioned 3D Head Generation with Self Score Distillation
Liu, Hongyu
Wang, Xuan
Wan, Ziyu
Shen, Yujun
Song, Yibing
Liao, Jing
Chen, Qifeng
Computer Vision and Pattern Recognition
This work presents HeadArtist for 3D head generation from text descriptions. With a landmark-guided ControlNet serving as the generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the supervision of the prior distillation itself. We call such a process self score distillation (SSD). In detail, given a sampled camera pose, we first render an image and its corresponding landmarks from the head model, and add some particular level of noise onto the image. The noisy image, landmarks, and text condition are then fed into the frozen ControlNet twice for noise prediction. Two different classifier-free guidance (CFG) weights are applied during these two predictions, and the prediction difference offers a direction on how the rendered image can better match the text of interest. Experimental results suggest that our approach delivers high-quality 3D head sculptures with adequate geometry and photorealistic appearance, significantly outperforming state-ofthe-art methods. We also show that the same pipeline well supports editing the generated heads, including both geometry deformation and appearance change.
title HeadArtist: Text-conditioned 3D Head Generation with Self Score Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.07539